Machine Learning Scientist

Posted Jul 28

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Create, train, and assess speech synthesis models, both autoregressive and non-autoregressive.

β€’ Initiate research into full-duplex and half-duplex multi-modal architectures, including unified S2S systems.

β€’ Select and refine speech representations: neural codecs, semantic tokens, mel features, and continuous latents.

β€’ Establish rigorous evaluations, both objective and perceptual, maintaining high standards for quality and prosodic control.

β€’ Partner with our linguists on TTS frontend behavior to ensure that modeling and frontend choices work in harmony.


⛳️ Requirements

β€’ Extensive knowledge of the speech synthesis literature, both contemporary and historical β€” including Tacotron, FastSpeech, VITS, VALL-E, and the codec-LM lineage. Possess insights on what has been effective and why.

β€’ Practical experience with neural codecs (EnCodec, DAC, Mimi, etc.) and various representation options.

β€’ Proficient in full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S).

β€’ Exceptional attention to detail regarding data quality, able to identify when an annotation pipeline is deteriorating or when an evaluation set is compromised.

β€’ Eager to engage in hands-on data and training tasks, with the initiative to develop pipelines to alleviate manual processes for the team.

β€’ Familiarity with TTS frontend components (G2P, normalization, prosody) and experience collaborating with linguists.

β€’ Strong foundational knowledge of PyTorch, comfortable with training loops, distributed training, and model internals.

β€’ PhD or equivalent research experience in speech, audio, ML, or computational linguistics, or a proven track record that renders formal credentials unnecessary.

β€’ Experience in multilingual TTS.

β€’ Background in prosody or paralinguistics.

β€’ Published research in speech, audio, or core ML venues.

β€’ Experience transitioning research models to production, including quantization, distillation, and streaming inference.


🏝️ Benefits

β€’ Competitive base salary plus significant early-stage equity.

β€’ Remote-friendly work environment.

β€’ Visa sponsorship available.

β€’ Access to a proprietary, full-duplex, studio-quality conversational speech corpus.

β€’ Provision of computing resources and tools necessary for the work.

β€’ Opportunity to have a direct impact on the future of voice AI.

People also viewed

Sourcegraph1 day ago

ML Engineer, Agentic Systems

North AmericaFull-timeMachine Learning Engineer$88k – $176k/year
ApplyView job
Quora1 day ago

Senior Machine Learning Engineer, Ads

US flagUnited States OnlyFull-timeMachine Learning Engineer$189.5k – $274.6k/year
ApplyView job
NBCUniversal1 day ago

Staff MLOps Engineer

CA flagCanada OnlyFull-timeMachine Learning Engineer
ApplyView job
Spotify1 day ago

Staff Machine Learning Engineer – Home Surfaces

US flagNew York OnlyFull-timeMachine Learning Engineer
ApplyView job
Fullscript1 day ago

Senior Machine Learning Engineer

CA flagCanada OnlyFull-timeMachine Learning EngineerC$140k – C$160k/year
ApplyView job
Vida Health1 day ago

Principal AI/ML Engineering Lead

US flagUnited States OnlyFull-timeMachine Learning Engineer$250k – $275k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers