Remotery

Machine Learning Engineer, Reliability

Posted Jul 28

This is a fully remote position, open to applicants in India, +2 more states.

📋 Description

• Take charge of the availability, latency, and throughput SLOs for a vast array of generative media model APIs that handle production traffic at scale.

• Develop the necessary monitoring, alerting, and observability to identify ML-specific failures, output quality issues, pipeline disruptions, and model regressions before they impact customers.

• Strengthen model deployment processes with canary releases, shadow testing, automated rollbacks, and validation gates to ensure the safe release of new model versions.

• Enhance the security framework of the model fleet, focusing on secure model serving, detection of abuse and misuse, rate limiting, and safeguarding against adversarial usage patterns.

• Implement operational safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that function reliably during inference without sacrificing performance.

• Lead incident response efforts for model API outages and performance degradations, conduct postmortems, and spearhead engineering initiatives to prevent future occurrences.

• Optimize capacity planning, autoscaling, and GPU fleet efficiency for inference workloads amid highly fluctuating traffic.

• Collaborate with model and infrastructure teams to integrate reliability, security, and safety requirements into the onboarding process of new models onto the platform.


⛳️ Requirements

• Minimum of 3 years of professional experience, with at least 1 year managing production ML or high-scale API systems, preferably with on-call responsibilities.

• Solid understanding of systems fundamentals, including distributed systems, networking, observability, and incident management.

• Familiarity with contemporary generative models (such as diffusion and transformers) and their failure modes in a production environment.

• Knowledge of security and safety practices for ML systems, experience in abuse prevention, content safety, or trust & safety engineering is highly advantageous.

• A strong inclination towards automation, measurement, and conducting blameless postmortems.


🏝️ Benefits

• You will have access to our extensive GPU cluster for inference and evaluation tasks.

• Some of the core technologies we utilize include Python, torch, diffusers, Kubernetes, and the fal Python SDK.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers