Join Rime as a Machine Learning Scientist to advance state-of-the-art speech synthesis and understanding models for enterprise voice AI. You will design, train, and evaluate cutting-edge autoregressive and non-autoregressive speech models, lead research on multi-modal architectures, and optimize speech representations. Requires deep expertise in speech synthesis literature, neural codecs, and PyTorch with a PhD or equivalent research experience.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
About The Company
Rime is a pioneering company specializing in voice AI technology designed for enterprise-level customer experience solutions. Our core expertise lies in developing high-volume, conversational text-to-speech (TTS) models that are engineered for accuracy, low latency, and flexible deployment in production environments. Unlike traditional voice AI providers, Rime emphasizes the importance of data quality and proprietary speech corpora, which serve as our competitive advantage. Our full-duplex, studio-quality conversational speech datasets, meticulously recorded and annotated by PhD linguists, form the foundation of our innovative models. Backed by top-tier investors including Unusual Ventures, Rime is at the forefront of research and development in speech synthesis and understanding, combining product innovation with rigorous scientific research to master the art of voice AI.
About The Role
We are seeking a talented and motivated Machine Learning Scientist to join our team. In this role, you will be instrumental in advancing the state-of-the-art in speech synthesis and speech understanding. Your primary responsibilities will include designing, training, and evaluating cutting-edge speech models, including autoregressive and non-autoregressive architectures. You will lead research initiatives on full-duplex and half-duplex multi-modal systems, exploring innovative architectures such as unified sequence-to-sequence models. Collaborating closely with linguists and engineers, you will refine speech representations, including neural codecs, semantic tokens, and mel features, ensuring the highest quality and prosodic control. Your work will directly impact the development of scalable, high-quality voice AI systems that meet the demanding needs of enterprise clients.
Qualifications
- Deep familiarity with speech synthesis literature, including models like Tacotron, FastSpeech, VITS, VALL-E, and codec-LM lineage
- Hands-on experience with neural codecs such as EnCodec, DAC, Mimi, and related representation techniques
- Experience with full- or half-duplex multi-modal modeling, including architectures like Moshi, LLaMA-Omni, or streaming sequence-to-sequence systems
- Strong attention to data quality, with the ability to identify issues in annotation pipelines and evaluation sets
- Willingness to handle data and training tasks that are unglamorous but essential, including building and maintaining pipelines
- Working knowledge of TTS frontend components such as G2P, normalization, and prosody, with experience collaborating with linguists
- Proficiency in PyTorch, including training loops, distributed training, and understanding of model internals
- PhD or equivalent research experience in speech, audio, machine learning, or computational linguistics, or a track record demonstrating similar expertise
Searching for Machine Learning & AI roles that provide visa sponsorship? Connect with international employers through Machine Learning & AI Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
- Design, train, and evaluate speech synthesis models, both autoregressive and non-autoregressive
- Lead research on multi-modal architectures, including full-duplex and half-duplex systems, to enhance speech understanding and synthesis
- Experiment with and optimize speech representations such as neural codecs, semantic tokens, and mel features to improve model performance
- Develop rigorous evaluation protocols, combining objective metrics and perceptual assessments to ensure high-quality output
- Collaborate with linguists to align modeling techniques with frontend speech processing components, ensuring seamless integration
- Build and maintain data pipelines, annotation processes, and training workflows to support ongoing research and production needs
- Stay updated with the latest advancements in speech synthesis and incorporate best practices into Rime’s models
- Contribute to the publication of research findings in core ML, speech, and audio venues, advancing the field
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
- Competitive salary package complemented by meaningful early-stage equity options
- Remote-friendly work environment with flexible scheduling
- Visa sponsorship available for international candidates
- Access to proprietary, full-duplex, studio-quality conversational speech corpus for research and development
- State-of-the-art compute resources and tooling to support advanced research and model training
- Opportunity to influence the future of voice AI technology and enterprise applications
- Collaborative environment with direct interaction with founders and industry experts
- High ownership culture with high standards and low bureaucracy
Interested in opportunities specifically in United State? Discover our dedicated Visa Sponsorship Jobs in United State page featuring roles from top employers in this location.
Rime is committed to fostering an inclusive and diverse workplace. We provide equal employment opportunities to all applicants and employees regardless of race, ethnicity, gender, sexual orientation, age, disability, or background. We believe that diverse perspectives drive innovation and excellence, and we strive to create an environment where everyone can thrive and contribute to our shared success.
Similar Jobs
Explore other opportunities that match your interests