
Machine Learning Engineer – Voice AI, Generative Music
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Mexico.
• Refine and enhance a singing-voice conversion and cloning model to achieve a minimum of 90% quality in blind listening evaluations.
• Accurately reproduce expressive vocal traits such as vibrato, falsetto, dynamics, and both spoken and sung delivery.
• Assess the expansion of the training dataset, alternative base models, and enterprise APIs that offer no-training guarantees.
• Take ownership of and enhance a Spanish-speaking Chatterbox Multilingual LoRA fine-tuning project.
• Guarantee precise reproduction of the Mexican accent and accurate pronunciation of J, Ñ, and X.
• Containerize the text-to-speech model and deploy it as a serverless inference endpoint utilizing RunPod or a comparable platform.
• Connect the inference endpoint to the existing web platform via an API.
• Train a second text-to-speech version using high-quality studio recordings.
• Create a proprietary Spanish-language lyrics generation model utilizing an open-weight LLM, along with LoRA fine-tuning, DPO, and RAG.
• Set evaluation standards and coordinate quality assessments with native Spanish speakers.
• Incorporate the finalized lyrics model into the current frontend.
• Minimum of 3 years of experience in training and deploying deep learning models in production, ideally in audio or NLP.
• Practical experience in at least two of the following: voice cloning or Singing Voice Conversion (SVC); fine-tuning Text-to-Speech models; or LLM fine-tuning using LoRA, DPO, and RAG.
• Strong command of PyTorch.
• Hands-on experience with GPU cloud infrastructure such as RunPod or AWS.
• Familiarity with Docker and serverless inference.
• Proven model evaluation skills, including benchmarks, ablation studies, and blind testing.
• Native or strong professional proficiency in Spanish, essential for lyrics evaluation and voice quality assurance.
• Preferred: background in music or experience with stems, MIDI, and DAWs.
• Preferred: experience in singing-voice synthesis with ACE Studio, ACE-Step, RVC, so-vits-svc, or similar technologies.
• Preferred: experience collaborating with licensed celebrity or artist voices and consent-based voice AI.
• Security through client vetting to reduce risks and guarantee reliable, timely payments.
• Career support and guidance in finding new opportunities if a project isn't the right match.
• Legal assistance regarding independent contractor or sole proprietorship status, taxes, and related processes.
• English language courses.
• Opportunities for professional growth.
• Team-building activities.
• Flexible working hours.
• 29 days of paid time off (18 working days per year plus all national holidays).
• 10 paid recovery days.
• Comprehensive financial and legal support for independent contractors.
• Complimentary English classes with native speakers or Ukrainian instructors.
• Dedicated HR support.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.