
Research Staff, Voice AI Foundations
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in Netherlands.
• Lead the development of Latent Space Models to tackle essential data, scaling, and cost challenges in robust, contextual voice AI.
• Create next-generation neural audio codecs that achieve extremely low-bit-rate compression while ensuring high-fidelity reconstruction.
• Develop steerable generative models capable of synthesizing a variety of human speech, including emotional tones, multi-speaker interactions, noisy environments, and overlapping speech scenarios.
• Construct embedding systems that decompose latent spaces into dimensions of speaker, content, style, environment, and channel.
• Utilize latent recombination techniques to produce synthetic audio data at unprecedented scales.
• Train multimodal speech-to-speech systems to enable universal understanding and empathetic, human-like responses.
• Design model architectures, training strategies, and inference algorithms optimized for bare-metal hardware.
• Facilitate cost-efficient training on datasets encompassing billions of hours and enable real-time inference for millions of simultaneous conversations.
• Perform rigorous controlled experiments, ablation studies, evaluations, and stress tests.
• Collaborate through research publications and contribute to open-source initiatives.
• Strong mathematical background in statistical learning theory, especially in self-supervised and multimodal learning contexts.
• Extensive expertise in foundational model architectures and scaling training processes across diverse modalities.
• Capacity to derive innovative mathematical formulations and implement them effectively.
• Proven ability to create and maintain extensive, high-quality, and diverse data pipelines.
• Experience in designing controlled experiments to isolate the impacts of architecture and validate theoretical insights.
• Proficiency in optimizing models for practical deployment, considering hardware constraints and efficiency strategies.
• A history of open-source contributions or research publications that advance the field of speech/language AI.
• Ability to quickly identify critical experiments that validate or disprove concepts.
• Vision to scale successful proofs-of-concept by a factor of 100.
• Strong mindset toward AI automation and augmentation.
• Scholarly literature or academic publications relevant to speech processing, STT/TTS, or similar fields requested in the application process.
• Legal authorization to work in the country where the position is based.
• Capability to meet visa sponsorship requirements for the location of the role.
• Remote work arrangement.
• AI-first work environment with opportunities to utilize and experiment with cutting-edge AI tools.
• Chance to engage in transformative voice AI research.
• Opportunity to contribute to open-source projects and research publications.
• AI Notetaker interview recording is optional; choosing to opt-out will not affect candidacy.
Deepgram
Deepgram
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.