
Staff MLOps Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in California, +1 more state.
• Construct ML infrastructure: Design, manage, and sustain resilient systems for low-latency model deployment, distributed inference pipelines, and automated real-time telemetry.
• Scale ranking systems: Transition models smoothly from experimentation to production, optimizing essential trade-offs between execution latency, GPU/CPU throughput, and cloud infrastructure expenses.
• Implement model CI/CD: Create dependable infrastructure for automated model versioning, canary releases, hot-swappable container deployments, and zero-downtime rollbacks.
• Enhance system observability: Design and monitor real-time pipelines to assess model performance, data distribution drift, and anomalies in system reliability.
• Develop evaluation loops: Engineer solid evaluation pipelines and feedback mechanisms to continuously verify live inference accuracy and avoid training-serving discrepancies.
• Optimize platform bottlenecks: Proactively identify and resolve performance bottlenecks throughout our serving layers, enhancing core tools, model warm-up durations, and researcher efficiency.
• Collaborate with research: Work closely with our internal ML researchers and backend engineers to convert experimental model innovations into robust, production-ready serving architectures.
• 4–8+ years of hands-on experience in MLOps, Machine Learning Engineering, or distributed platform/infrastructure engineering.
• Proven experience in deploying and serving ultra-low-latency machine learning models under significant, real-time concurrent workloads.
• Maintain extensive, production-grade expertise in Python and PyTorch.
• Operate proficiently across major cloud platforms (AWS, GCP, or Azure) using contemporary containerization and orchestration tools (Docker, Kubernetes).
• Demonstrate experience in designing robust, scalable data pipelines, model registries (e.g., MLflow), and automated CI/CD infrastructures.
• Possess a solid, foundational understanding of the complete machine learning lifecycle, asynchronous event-driven patterns, and distributed systems.
• Comprehensive premium medical/dental/vision coverage
• Unlimited paid time off
• Highly collaborative, world-class engineering culture
Sword Health
Apiux Tech
Lime
May Mobility
Get handpicked remote jobs straight to your inbox weekly.