
Machine Learning Engineer – Voice Conversion
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in United States.
• Architect, implement, pre-train, fine-tune, and post-train/alignment of large-scale speech models.
• Design, execute, and evaluate scientific experiments to enhance understanding of models.
• Develop and refine development tools to boost team productivity.
• Contribute throughout the stack, from low-level optimizations to high-level model architecture.
• Define data requirements and work collaboratively on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
• Create automated objective and subjective evaluations, which include listening tests, SV/WER/ASR-based metrics, robustness and bias assessments, and red-team studies.
• Strengthen the training-to-evaluation-to-inference pipeline; assess latency, memory, and costs; and ensure compliance with production SLAs through monitoring and rollback.
• Assist in establishing safety and consent measures as well as mitigating misuse and abuse for responsible speech technology.
• Collaborate with research, data, infrastructure, and product teams to deliver measurable enhancements to speech systems.
• Outstanding research and development experience with large-scale audio models (over 8B parameters, exceeding 500k hours of data).
• Extensive hands-on experience with diffusion and/or flow-matching transformers, encompassing samplers, schedules, conditioning mechanisms, and distillation.
• Profound hands-on experience in training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
• Significant experience with multi-node, multi-GPU distributed training utilizing FSDP, DeepSpeed, or similar technologies.
• Strong software engineering abilities with a solid history of developing complex systems.
• High proficiency in PyTorch and performance optimization, including profiling and CUDA/Triton/C++ as required.
• Experience in writing reliable, production-quality code.
• Proven track record of deploying large-scale speech/audio or multimodal generative models in production.
• Background in handling large-scale ML data and assessing quality through both subjective and objective signals.
• Experience with voice cloning, speech control/steerability, or expressive speech generation.
• Noteworthy publications and/or contributions to open-source projects in speech, audio, or machine learning.
• Competitive salary and substantial company equity.
• Medical, dental, and vision insurance – with 99.99% of premiums covered by Cantina.
• 42 days of paid time off, which includes 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays.
• Generous parental leave and fertility support.
• 401(k) retirement savings plan.
• Lifestyle spending account – $500/month for personal use.
• Complimentary lunch and snacks for employees in the office.
• One Medical membership, among other benefits!
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.