
Machine Learning Engineer, Speech – Joint Audio-Video Modeling
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Europe.
• Audio Representations: Develop, train, and enhance the audio VAEs, neural codecs, and vocoders that our generative models utilize, focusing on latent design, reconstruction, perceptual objectives, and the balance between compression and fidelity.
• Model Building: Design, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for extensive audio and video generation.
• Joint Audio-Video Modeling: Create audio conditioning and cross-modal alignment within joint AV models, incorporating audio latents with video latents, reference audio, multi-speaker conditioning, and multi-shot generation for audio/video modeling.
• Experimental Design: Design, execute, and assess scientific experiments to enhance our understanding of the models.
• Data Ownership: Establish data requirements and collaborate on acquisition, curation, AV synchronization, quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.
• Rigorous Evaluation: Create automated objective and subjective evaluations for audio fidelity, intelligibility metrics, AV synchronization, listening and viewing tests, robustness and bias assessments, and red-team studies.
• Inference Efficiency: Lead efforts in distillation, step-count reduction, quantization, and kernel/memory optimization to achieve interactive latency and cost objectives.
• Pipeline Delivery: Strengthen the training → evaluation → inference pipeline; monitor latency, memory, and costs; and ensure adherence to production service level agreements (SLAs) with effective monitoring and rollback mechanisms.
• GPU Scaling: Collaborate with infrastructure teams to execute distributed training/inference on cloud platforms and ensure the reliability and observability of productionized models.
• Project Leadership: Independently steer small research projects while contributing to larger team initiatives, including collaborative efforts on video generation.
• Tool Development: Create and enhance development tools to boost team productivity.
• Safety & Responsibility: Engage in establishing safety and consent guardrails, watermarking, and measures to mitigate misuse or abuse of voice and likeness technology.
• Outstanding research and development experience with large-scale audio models (>8B parameters, >500k hours of data).
• Extensive hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
• Significant hands-on experience in training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
• Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
• Solid software engineering skills with a proven history of building complex systems.
• Proficiency in PyTorch and performance optimization (profiling, CUDA/Triton/C++ as necessary) along with the ability to write reliable, production-quality code.
• Experience in deploying large-scale speech/audio or multimodal generative models into production.
• Background in managing large-scale machine learning data, with the capacity to iterate on data and evaluate quality using both subjective and objective metrics.
• Familiarity with voice cloning, speech control/steerability, or expressive speech generation.
• Recognized publications and/or contributions to open-source projects in the fields of speech/audio/ML.
• Strongly preferred:
• Experience in multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes that include dialogue, music, and sound design generated in conjunction with video, as well as the necessary cross-modal alignment to maintain synchronization.
• Knowledge in video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, and constructing data pipelines for video models.
• Expertise in streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).
• Competitive salary and generous company equity.
• Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina.
• 42 days of paid time off, including:
• 15 PTO days.
• 10 sick days.
• 15 company holidays.
• 2 floating holidays.
• Generous parental leave and fertility support.
• 401(k) retirement savings plan.
• Lifestyle spending account – $500/month to use as you wish.
• Complimentary lunch and snacks for in-office employees.
• One Medical membership, and more!
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.