Remotery

Machine Learning Engineer, Speech – Joint Audio-Video Modeling

atCantinaRemoteEuropeFull-timeMachine Learning EngineerMid-levelSenior$200k – $220k/year

Posted Jul 30

This is a fully remote position, open to applicants in Europe.

📋 Description

• Audio Representations: Develop, train, and enhance the audio VAEs, neural codecs, and vocoders that our generative models utilize, focusing on latent design, reconstruction, perceptual objectives, and the balance between compression and fidelity.

• Model Building: Design, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for extensive audio and video generation.

• Joint Audio-Video Modeling: Create audio conditioning and cross-modal alignment within joint AV models, incorporating audio latents with video latents, reference audio, multi-speaker conditioning, and multi-shot generation for audio/video modeling.

• Experimental Design: Design, execute, and assess scientific experiments to enhance our understanding of the models.

• Data Ownership: Establish data requirements and collaborate on acquisition, curation, AV synchronization, quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.

• Rigorous Evaluation: Create automated objective and subjective evaluations for audio fidelity, intelligibility metrics, AV synchronization, listening and viewing tests, robustness and bias assessments, and red-team studies.

• Inference Efficiency: Lead efforts in distillation, step-count reduction, quantization, and kernel/memory optimization to achieve interactive latency and cost objectives.

• Pipeline Delivery: Strengthen the training → evaluation → inference pipeline; monitor latency, memory, and costs; and ensure adherence to production service level agreements (SLAs) with effective monitoring and rollback mechanisms.

• GPU Scaling: Collaborate with infrastructure teams to execute distributed training/inference on cloud platforms and ensure the reliability and observability of productionized models.

• Project Leadership: Independently steer small research projects while contributing to larger team initiatives, including collaborative efforts on video generation.

• Tool Development: Create and enhance development tools to boost team productivity.

• Safety & Responsibility: Engage in establishing safety and consent guardrails, watermarking, and measures to mitigate misuse or abuse of voice and likeness technology.


⛳️ Requirements

• Outstanding research and development experience with large-scale audio models (>8B parameters, >500k hours of data).

• Extensive hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.

• Significant hands-on experience in training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.

• Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).

• Solid software engineering skills with a proven history of building complex systems.

• Proficiency in PyTorch and performance optimization (profiling, CUDA/Triton/C++ as necessary) along with the ability to write reliable, production-quality code.

• Experience in deploying large-scale speech/audio or multimodal generative models into production.

• Background in managing large-scale machine learning data, with the capacity to iterate on data and evaluate quality using both subjective and objective metrics.

• Familiarity with voice cloning, speech control/steerability, or expressive speech generation.

• Recognized publications and/or contributions to open-source projects in the fields of speech/audio/ML.

• Strongly preferred:

• Experience in multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes that include dialogue, music, and sound design generated in conjunction with video, as well as the necessary cross-modal alignment to maintain synchronization.

• Knowledge in video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, and constructing data pipelines for video models.

• Expertise in streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).


🏝️ Benefits

• Competitive salary and generous company equity.

• Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina.

• 42 days of paid time off, including:

• 15 PTO days.

• 10 sick days.

• 15 company holidays.

• 2 floating holidays.

• Generous parental leave and fertility support.

• 401(k) retirement savings plan.

• Lifestyle spending account – $500/month to use as you wish.

• Complimentary lunch and snacks for in-office employees.

• One Medical membership, and more!

People also viewed

Doma2 days ago

Staff Machine Learning Engineer

US flagCalifornia OnlyFull-timeMachine Learning Engineer$165.2k – $236.3k/year
ApplyView job
CSC Generation2 days ago

Senior Machine Learning Engineer, Causal & Decision Systems

US flagTexas OnlyFull-timeMachine Learning Engineer
ApplyView job
Accelerant2 days ago

Principal Machine Learning Engineer

US flagUnited States, +1 more stateFull-timeMachine Learning Engineer
ApplyView job
Capgemini2 days ago

Senior Machine Learning Engineer

CO flagColombia OnlyFull-timeMachine Learning Engineer
ApplyView job
Airbnb2 days ago

Senior Machine Learning Engineer, Trust

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $235k/year
ApplyView job
CrowdStrike2 days ago

Senior Machine Learning Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer$140k – $215k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers