
Research Staff – Voice AI Foundations
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in Australia.
• Create cutting-edge neural audio codecs featuring ultra-low-bit-rate compression and high-fidelity reconstruction.
• Lead the development of steerable generative models for a variety of human speech scenarios, including emotional tones, multiple speakers, background noise, and overlapping speech.
• Construct embedding systems that decompose latent representations into dimensions of speaker, content, style, environment, and channel.
• Utilize latent recombination techniques to produce synthetic audio data on scales previously deemed unfeasible.
• Train multimodal speech-to-speech systems to ensure robust comprehension and empathetic, human-like interactions.
• Design model architectures, training methodologies, and inference algorithms optimized for bare-metal hardware performance.
• Facilitate cost-effective training on datasets encompassing billions of hours, enabling real-time inference for millions of simultaneous conversations.
• Execute thorough controlled experiments, ablation studies, evaluations, stress testing, and analyses of edge cases.
• Contribute to pioneering research that advances the fields of voice AI, speech, and language modeling.
• Solid mathematical grounding in statistical learning theory, especially in the areas of self-supervised and multimodal learning.
• Profound knowledge of foundational model architectures and the ability to scale training across various modalities.
• Capacity to derive innovative mathematical formulations and implement them in an efficient manner.
• Proven experience in constructing and curating extensive datasets while ensuring quality and diversity.
• Demonstrated history of designing controlled experiments that validate architectural innovations and theoretical concepts.
• Experience in optimizing models for deployment in real-world settings, taking into account hardware limitations and efficiency strategies.
• Background in open-source contributions or research publications that advance the fields of speech and language AI.
• Comfort in actively using and experimenting with advanced AI tools.
• Ability to quickly identify critical experiments that can validate or refute ideas.
• Vision to scale successful proofs-of-concept by a factor of 100.
• Familiarity with neural audio codecs, generative models, embedding systems, latent representations, multimodal speech-to-speech systems, model architectures, training methodologies, inference algorithms, and efficient hardware deployment.
• Flexible remote work options.
• An AI-driven work environment that promotes the use and experimentation with advanced AI tools.
• An opportunity to engage in transformative voice AI research at scale.
• Opportunities for open-source contributions and research publications.
• The option to decline AI Notetaker interview recordings without impacting candidacy.
Deepgram
Deepgram
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.