
Senior Machine Learning Engineer, Voice Agents
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in France.
• Take charge of the architecture for significant components of the speech-to-speech open-source library, focusing on pipeline design, latency management, and real-time loop reliability.
• Implement new ASR, TTS, and end-to-end speech models while ensuring clean abstractions are preserved.
• Evaluate community pull requests, manage issues, release updates, and foster contributions from project collaborators.
• Develop the hf-voice developer API and streaming protocol, which encompasses session lifecycle, WebSockets/WebRTC transport, authentication, error semantics, and versioning.
• Create real-time GPU inference serving that includes concurrency, autoscaling, observability, and optimization of cost-per-session.
• Partner with Hub and inference teams to streamline the integration of voice agents into products and demonstrations.
• Transition hf-voice from a demonstration phase to production readiness through load testing, establishing SLOs, and ensuring graceful degradation.
• Produce documentation, examples, and templates for developers.
• Provide support for existing deployments, starting with the Reachy Mini fleet.
• Engage in public discussions about the work through blog posts, demonstrations, or conference presentations.
• Senior engineer capable of autonomously owning and advancing a significant portion of an architecture.
• Proven experience in building developer-facing infrastructure in an AI or developer-tools company, including inference APIs or agent infrastructure.
• Significant open-source contributions to a Python library.
• Proficient in async Python and distributed systems, including an understanding of their potential failure modes.
• Experience in deploying real-time systems involving streaming, WebSockets, WebRTC, audio or video pipelines, or live inference.
• Practical production experience with LLMs or multimodal models.
• Strong written communication skills and the ability to collaborate asynchronously and publicly.
• Passion for voice and conversational AI.
• Contributions to voice-agent frameworks such as speech-to-speech, Pipecat, LiveKit Agents, Vocode, or TEN.
• Contributions to llama.cpp or another low-level inference runtime.
• Hands-on experience with ASR, TTS, or end-to-end speech models, including evaluating latency and quality trade-offs.
• Experience with GPU serving, quantization, or on-device inference.
• Knowledge of audio pipelines, including VAD, echo cancellation, jitter buffers, barge-in, and turn detection.
• Experience delivering solutions for embedded or robotics targets.
• A public track record demonstrated through talks, blog posts, or demos.
• A workplace committed to diversity, equity, and inclusivity.
• An equal opportunity employer with a strong commitment to nondiscrimination.
• Reimbursement for relevant conferences, training, and educational opportunities.
• Flexible work hours.
• Options for remote work.
• Comprehensive health, dental, and vision benefits for employees and their dependents.
• Parental leave.
• Flexible paid time off.
• Opportunity for remote employees to visit offices in NYC and Paris.
• Workstation equipment provided as required.
• Equity participation for all employees.
• A community that actively supports the ML/AI community.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.