
Director, Text-to-Speech Synthesis Research
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in California, +1 more state.
• Take ownership of the TTS research and model roadmap.
• Determine which technical strategies can significantly enhance speech-generation quality.
• Propel advancements in neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategies, as well as post-training and inference performance.
• Ensure that research developments lead to quantifiable improvements in production.
• Evaluate research, question assumptions, design experiments, troubleshoot model failures, and address high-impact technical challenges.
• Develop evaluation and benchmarking processes utilizing automated metrics and human perceptual assessments.
• Supervise individual contributors and technical lead managers.
• Recruit and nurture researchers and research leaders while upholding a high technical standard.
• Mentor senior researchers to become technical leaders and define direction across sub-teams.
• Collaborate with engineering and product leadership to ensure ship-readiness.
• Represent Deepgram’s TTS research both internally and externally.
• Extensive expertise in contemporary TTS, speech generation, or audio generative modeling.
• Proven history of personally training and enhancing large-scale neural models.
• Proficient understanding of the modern speech-generation stack and the existing challenges regarding naturalness, expressiveness, controllability, robustness, voice consistency, and inference costs.
• Experience in setting research direction amidst uncertainty, prioritizing experiments, allocating compute resources and researcher time, and discontinuing ineffective methods.
• Proven experience in leading researchers and research engineers through other technical leaders, developing tech lead managers or their equivalents, and defining direction across sub-teams.
• AI as the fundamental mode of operation, with experience restructuring workflows around AI and a clear perspective on its limitations in speech research.
• Capability to articulate complex technical trade-offs to product, engineering, and executive audiences.
• Scholarly publications or academic literature related to speech processing, STT/TTS, or similar fields.
• Technical proficiency with technologies, systems, or tools used for training AI models, including PyTorch or other relevant libraries.
• Legally authorized to work in the United States; visa sponsorship requirements must be disclosed.
• No specific educational credentials required.
• Equity.
• Bonus.
• Base salary compensation.
• AI Notetaker interview recording/transcription, with the option to opt out without affecting candidacy.
Mercy Health
Mercy Health
Bon Secours
Bon Secours
Get handpicked remote jobs straight to your inbox weekly.