I

Principal Research Scientist - Speech and Audio Foundation Models

intelix.ai San Francisco Bay Area
Visa Sponsorship Relocation
Apply
AI Summary

We are seeking a Principal Research Scientist to lead the development of speech and audio foundation models. The successful candidate will be responsible for building and improving voice and speech models, designing evaluation frameworks, and taking models into production. The ideal candidate will have hands-on experience with foundation model training and real voice or speech research.

Key Highlights
Build and improve voice and speech models across TTS, STT, and speech-to-speech
Design evaluation frameworks to prove the work, including benchmarks and failure analysis
Take models into production alongside the serving engineering team
Key Responsibilities
Train foundation models, including pre-training, reinforcement learning, and post-training
Build and improve voice and speech models across TTS, STT, and speech-to-speech
Design evaluation frameworks to prove the work, including benchmarks and failure analysis
Take models into production alongside the serving engineering team
Technical Skills Required
Machine Learning Natural Language Processing Speech Synthesis
Benefits & Perks
Base salary: $270,000-$500,000
Bonus and equity
Benefits
Relocation assistance
Visa transfer support
Nice to Have
Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech, or ICASSP
PhD in ML or NLP, or equivalent practical experience
Frontier exposure: multimodal, agents, tool use, test-time compute
Public work: side projects, open-source, technical write-ups

Job Description


Principal Research Scientist Speech & Audio Foundation Models

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.


$270,000–$500,000 base plus bonus, equity and benefits (US).

Relocation assistance available. Visa transfer supported

San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.

Permanent, full-time.



A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.


Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.



The role


Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.



Responsibilities



  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.



Essential



  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.



Desirable



  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Frontier exposure: multimodal, agents, tool use, test-time compute.
  • Public work: side projects, open-source, technical write-ups.


Similar Jobs

Explore other opportunities that match your interests

Founding Senior Backend Engineer

Programming
59m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Stealth Startup

San Francisco Bay Area

Full Stack Engineer

Programming
5h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

coffeespace

San Francisco Bay Area

Staff Engineer, Safety Experience (Trust & Safety Systems)

Programming
17h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Discord

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly