
Machine Learning Engineer – Voice AI, Generative Music
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Refine and enhance a singing voice conversion and cloning model.
• Elevate vocal fidelity from a baseline of 70–75% to a minimum of 90% in blind listening assessments.
• Accurately reproduce expressive vocal traits, including vibrato, falsetto, dynamics, spoken delivery, and sung delivery.
• Assess the expansion of the training dataset, explore alternative base models, and evaluate enterprise APIs.
• Take ownership of and enhance a validated Chatterbox Multilingual LoRA fine-tune tailored for a Spanish-speaking voice.
• Guarantee precise reproduction of the Mexican accent and correct pronunciation of J, Ñ, and X.
• Containerize the text-to-speech model and deploy it as a serverless inference endpoint utilizing RunPod or a comparable service.
• Connect the inference endpoint to the current web platform via an API.
• Train a secondary version using clean studio recordings to enhance output consistency.
• Create a proprietary model for generating Spanish-language lyrics.
• Leverage an open-weight LLM, apply LoRA fine-tuning, utilize DPO preference optimization, and implement RAG over a curated content corpus.
• Define evaluation standards and coordinate quality assessments with native Spanish speakers.
• Seamlessly integrate the finalized lyrics model into the existing frontend.
• Over 3 years of experience in training and deploying deep learning models in production, ideally in audio or NLP domains.
• Practical experience in at least two of the following: voice cloning or Singing Voice Conversion (SVC); fine-tuning Text-to-Speech models; LLM fine-tuning using LoRA, DPO, and RAG.
• Strong expertise in PyTorch.
• Hands-on experience with GPU cloud infrastructure such as RunPod or AWS.
• Familiarity with Docker and serverless inference methodologies.
• Thorough, evidence-based model evaluation practices, including benchmarks, ablation studies, and blind testing.
• Native or strong professional proficiency in Spanish, essential for lyrics evaluation and voice quality assurance.
• Background in music or experience with music production tools and workflows, including stems, MIDI, and DAWs (preferred).
• Experience with singing voice synthesis solutions like ACE Studio, ACE-Step, RVC, so-vits-svc, or similar technologies (preferred).
• Experience working with licensed celebrity or artist voices and consent-based voice AI (preferred).
• Security through client vetting to reduce risks and ensure reliable, timely payments.
• Career support and assistance in finding new opportunities.
• Legal assistance regarding independent contractor or sole proprietorship status, taxes, and related processes.
• English language courses.
• Opportunities for professional growth.
• Team-building events.
• Flexible working hours.
• 29 days of paid time off (18 working days per year plus all national holidays).
• 10 paid recovery days.
• Comprehensive financial and legal support for independent contractors.
• Free English classes with native speakers or Ukrainian instructors.
• Dedicated HR support.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.