
Research Staff, Voice AI Foundations
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in United Kingdom.
• Create cutting-edge neural audio codecs featuring extremely low-bit-rate compression and high-fidelity reconstruction across extensive global audio datasets.
• Lead the development of steerable generative models capable of synthesizing a wide range of human speech, including emotional nuances, multi-speaker situations, environmental sounds, and overlapping dialogues.
• Build embedding systems that decompose codec latent spaces into understandable dimensions such as speaker, content, style, environment, and channel.
• Employ latent recombination techniques to produce synthetic audio data on an unprecedented scale.
• Train multimodal speech-to-speech systems that comprehend diverse human inputs and generate empathetic, human-like responses.
• Design model architectures, training methodologies, and inference algorithms tailored for bare-metal hardware to ensure cost-effective training on billion-hour datasets and real-time inference.
• Conduct foundational research in Latent Space Models to tackle challenges related to data, scale, and cost in voice AI.
• Create controlled experiments, ablation studies, evaluations, stress tests, and benchmarks to validate research hypotheses.
• Collaborate through contributions to open-source projects and research publications aimed at advancing speech and language AI.
• Strong mathematical background in statistical learning theory, especially in areas pertinent to self-supervised and multimodal learning.
• Extensive expertise in foundational model architectures and scaling training across various modalities.
• Proven capability to bridge theoretical concepts and practical implementations by deriving innovative mathematical formulations and executing them effectively.
• Demonstrated experience in constructing data pipelines that process and curate large datasets while ensuring quality and diversity.
• A history of designing controlled experiments that isolate architectural innovations and confirm theoretical insights.
• Experience in optimizing models for real-world applications, including constraints related to hardware and efficiency strategies.
• A track record of open-source contributions or research publications that advance the field of speech and language AI.
• Ability to quickly identify critical experiments that can validate or refute ideas.
• Vision for scaling successful proofs-of-concept by a factor of 100.
• Comfort in using AI to automate processes and enhance personal impact.
• Agility in adapting, experimenting, learning continuously, and thriving in a swiftly evolving AI landscape.
• Flexible remote work arrangement.
• An AI-first work environment that encourages active use and experimentation with advanced AI tools.
• Opportunity to lead pioneering foundational research in voice AI with transformative implications.
• Chance to contribute to open-source initiatives and research publications.
• Optional AI Notetaker interview recording; choosing not to participate does not affect candidacy.
Deepgram
Deepgram
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.