
Senior Machine Learning Engineer, ML Infrastructure – Online
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Washington.
• Design and manage large-scale online inference infrastructure that serves production ML models with minimal latency and maximum reliability.
• Create infrastructure that supports distributed training workflows.
• Integrate ML pipelines with workflow orchestration systems to ensure reliable multi-stage training processes.
• Enhance model performance through various techniques including compilation, GPU/CPU utilization optimization, request scheduling, kernel fusion, and runtime tuning.
• Increase observability by monitoring latency, throughput, error rates, costs, saturation, and model health.
• Collaborate with ML engineers to accelerate model iterations while ensuring production safety, scalability, and cost-effectiveness.
• Enhance the reliability and reproducibility of model serving workflows, which includes packaging, artifact validation, compatibility testing, and automation of deployments.
• Drive architectural enhancements to make the online ML platform more robust, user-friendly, scalable, and cost-efficient.
• Proven experience in building and operating production-grade online ML inference systems.
• Familiarity with model serving frameworks like NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or equivalent systems.
• Experience in optimizing inference workloads through dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
• Strong background in distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
• Proficient programming skills in Python, with hands-on experience in production ML systems and high-scale services.
• Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and management of serving lifecycles.
• Experience in designing infrastructure for safe model rollouts, canary testing, A/B testing, and automated rollbacks.
• Strong systems thinking with the ability to analyze latency, throughput, reliability, scalability, and cost trade-offs.
• Demonstrated ability to lead technical direction and influence architectural decisions across teams without formal authority.
• Adequate proficiency in English for professional verbal and written communication.
• Work visa/immigration sponsorship is not available for this position.
• Relocation support is not available for this position.
• Comprehensive health, life, and disability insurance.
• Commute subsidy.
• Employee stock ownership.
• Competitive retirement and pension plans.
• Generous vacation and personal days.
• Support for new parents through leave and family-care programs.
• Office snacks available.
• Mental health and wellbeing programs and support.
• Employee Resource Groups.
• Global Employee Assistance Program.
• Training and development opportunities.
• Volunteering and donation matching initiatives.
• Equity awards.
• Eligibility for participation in company incentive plans, including annual discretionary bonuses or sales commissions.
SumerSports
Airbnb
The Home Depot
Get handpicked remote jobs straight to your inbox weekly.