
Machine Learning Engineer – Voice AI, Generative Music
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Spain.
• Take charge of the complete machine-learning workflow, encompassing audio data preparation, training experiments, optimized inference, and production deployment.
• Enhance and refine a singing-voice conversion and cloning model to reach a minimum of 90% fidelity in blind listening assessments.
• Recreate expressive vocal qualities such as vibrato, falsetto, dynamics, spoken delivery, and sung delivery.
• Assess dataset expansion, alternative base models, and enterprise APIs with no-training guarantees.
• Manage an existing Chatterbox Multilingual LoRA fine-tune tailored for a Spanish-speaking voice.
• Ensure precise reproduction of a Mexican accent and accurate pronunciation of J, Ñ, and X sounds.
• Containerize and deploy the text-to-speech model as a serverless inference endpoint using RunPod or a comparable platform.
• Connect the inference endpoint to the current web platform through an API.
• Train a secondary text-to-speech version utilizing clean studio recordings to enhance output stability.
• Create a proprietary Spanish-language lyrics-generation model leveraging an open-weight LLM, LoRA fine-tuning, DPO, and RAG.
• Set evaluation standards and coordinate quality assessments with native Spanish speakers.
• Incorporate the finalized lyrics model into the existing frontend.
• Over 3 years of experience in training and deploying deep learning models in production, ideally within audio or NLP domains.
• Hands-on expertise in at least two of the following: Voice cloning or Singing Voice Conversion (SVC); Text-to-Speech model fine-tuning; LLM fine-tuning using LoRA, DPO, and RAG.
• Strong command of PyTorch.
• Practical experience with GPU cloud infrastructures such as RunPod or AWS, Docker, and serverless inference.
• Evidence-based model evaluation skills, including benchmarks, ablation studies, and blind testing.
• Native or strong professional proficiency in Spanish, essential for lyrics evaluation and voice quality assurance.
• Preferred: Background in music or experience with stems, MIDI, and DAWs.
• Preferred: Familiarity with ACE Studio, ACE-Step, RVC, so-vits-svc, or similar singing-voice synthesis technologies.
• Preferred: Experience with licensed celebrity or artist voices and consent-based voice AI.
• Security through client vetting to reduce risks and ensure reliability and prompt payments.
• Career support and assistance in finding new opportunities if a project does not align well.
• Legal assistance regarding independent contractor or sole proprietorship status, taxes, and related matters.
• English language courses.
• Opportunities for professional growth.
• Team-building activities.
• Flexible working hours.
• 29 days of paid time off (18 working days per year plus all national holidays).
• 10 paid recovery days.
• Comprehensive financial and legal support for independent contractors.
• Complimentary English classes with native speakers or Ukrainian educators.
• Dedicated HR support.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.