
Machine Learning Scientist
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
β’ Create, train, and assess speech synthesis models, both autoregressive and non-autoregressive.
β’ Initiate research into full-duplex and half-duplex multi-modal architectures, including unified S2S systems.
β’ Select and refine speech representations: neural codecs, semantic tokens, mel features, and continuous latents.
β’ Establish rigorous evaluations, both objective and perceptual, maintaining high standards for quality and prosodic control.
β’ Partner with our linguists on TTS frontend behavior to ensure that modeling and frontend choices work in harmony.
β’ Extensive knowledge of the speech synthesis literature, both contemporary and historical β including Tacotron, FastSpeech, VITS, VALL-E, and the codec-LM lineage. Possess insights on what has been effective and why.
β’ Practical experience with neural codecs (EnCodec, DAC, Mimi, etc.) and various representation options.
β’ Proficient in full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S).
β’ Exceptional attention to detail regarding data quality, able to identify when an annotation pipeline is deteriorating or when an evaluation set is compromised.
β’ Eager to engage in hands-on data and training tasks, with the initiative to develop pipelines to alleviate manual processes for the team.
β’ Familiarity with TTS frontend components (G2P, normalization, prosody) and experience collaborating with linguists.
β’ Strong foundational knowledge of PyTorch, comfortable with training loops, distributed training, and model internals.
β’ PhD or equivalent research experience in speech, audio, ML, or computational linguistics, or a proven track record that renders formal credentials unnecessary.
β’ Experience in multilingual TTS.
β’ Background in prosody or paralinguistics.
β’ Published research in speech, audio, or core ML venues.
β’ Experience transitioning research models to production, including quantization, distillation, and streaming inference.
β’ Competitive base salary plus significant early-stage equity.
β’ Remote-friendly work environment.
β’ Visa sponsorship available.
β’ Access to a proprietary, full-duplex, studio-quality conversational speech corpus.
β’ Provision of computing resources and tools necessary for the work.
β’ Opportunity to have a direct impact on the future of voice AI.
Sourcegraph
Quora
NBCUniversal
Spotify
Get handpicked remote jobs straight to your inbox weekly.