
Director of Research, Text to Speech
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in California, +1 more state.
β’ Take charge of the TTS research and model roadmap, making technical decisions that significantly enhance speech-generation quality.
β’ Propel advancements in neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategies, post-training, and inference performance.
β’ Ensure that research progress translates into quantifiable production improvements.
β’ Analyze research, question assumptions, design experiments, troubleshoot model failures, and resolve high-impact technical challenges.
β’ Develop evaluation and benchmarking techniques that integrate automated metrics with human perceptual assessments.
β’ Lead individual contributors and technical lead managers.
β’ Recruit and nurture researchers and technical leaders, uphold a high technical standard, and guide direction across sub-teams.
β’ Collaborate with engineering and product leadership to ensure ship-readiness.
β’ Represent Deepgramβs TTS research both internally and externally.
β’ Establish the team and operational model for rapid, high-quality research execution.
β’ Profound expertise in contemporary TTS, speech generation, or audio generative modeling.
β’ Proven history of personally training and enhancing large-scale neural models.
β’ In-depth knowledge of the modern speech-generation stack and existing challenges regarding naturalness, expressiveness, controllability, robustness, voice consistency, and inference costs.
β’ Experience in directing research amidst uncertainties, prioritizing experiments, managing compute and researcher time, and discontinuing ineffective methodologies.
β’ Background in leading researchers and research engineers through other technical leaders, including the development of tech lead managers or equivalent roles.
β’ Familiarity with AI as a default operational mode, including experience in restructuring workflows around AI and understanding its limitations in speech research.
β’ Capability to convey complex technical trade-offs to product, engineering, and executive stakeholders.
β’ Experience with TTS or generative-audio models deployed at significant production scale.
β’ Proven track record in building or significantly scaling a high-performing AI research organization.
β’ Experience with evaluation systems for generative speech, expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
β’ Acknowledged external contributions such as publications, open-source projects, patents, or invited presentations.
β’ Experience in dynamic startup or research environments where models transition from concept to production.
β’ Must be legally authorized to work in the location where the position is based.
β’ Must address visa sponsorship needs in the application.
β’ Equity
β’ Bonus
β’ Remote work arrangement
β’ AI Notetaker interview recording and transcription, with the option to opt out without affecting candidacy.
Ogilvy Health
Ogilvy
Coopers Group AG
BioMarin Pharmaceutical Inc.
Get handpicked remote jobs straight to your inbox weekly.