
Machine Learning Engineer, Reliability
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in India, +2 more states.
• Take charge of the availability, latency, and throughput SLOs for a vast array of generative media model APIs that handle production traffic at scale.
• Develop the necessary monitoring, alerting, and observability to identify ML-specific failures, output quality issues, pipeline disruptions, and model regressions before they impact customers.
• Strengthen model deployment processes with canary releases, shadow testing, automated rollbacks, and validation gates to ensure the safe release of new model versions.
• Enhance the security framework of the model fleet, focusing on secure model serving, detection of abuse and misuse, rate limiting, and safeguarding against adversarial usage patterns.
• Implement operational safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that function reliably during inference without sacrificing performance.
• Lead incident response efforts for model API outages and performance degradations, conduct postmortems, and spearhead engineering initiatives to prevent future occurrences.
• Optimize capacity planning, autoscaling, and GPU fleet efficiency for inference workloads amid highly fluctuating traffic.
• Collaborate with model and infrastructure teams to integrate reliability, security, and safety requirements into the onboarding process of new models onto the platform.
• Minimum of 3 years of professional experience, with at least 1 year managing production ML or high-scale API systems, preferably with on-call responsibilities.
• Solid understanding of systems fundamentals, including distributed systems, networking, observability, and incident management.
• Familiarity with contemporary generative models (such as diffusion and transformers) and their failure modes in a production environment.
• Knowledge of security and safety practices for ML systems, experience in abuse prevention, content safety, or trust & safety engineering is highly advantageous.
• A strong inclination towards automation, measurement, and conducting blameless postmortems.
• You will have access to our extensive GPU cluster for inference and evaluation tasks.
• Some of the core technologies we utilize include Python, torch, diffusers, Kubernetes, and the fal Python SDK.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.