
Senior Machine Learning Engineer, ML Infrastructure
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Washington.
• Design and manage large-scale online inference infrastructure that supports production ML models with minimal latency and exceptional reliability.
• Develop infrastructure that facilitates distributed training workflows utilizing PyTorch, Ray Data, and Ray Train.
• Integrate ML pipelines with workflow orchestration platforms such as Flyte or Airflow.
• Enhance model performance through techniques such as model compilation, improvements in GPU/CPU utilization, request scheduling, kernel fusion, and runtime tuning.
• Improve observability of ML systems by monitoring latency, throughput, error rates, costs, saturation, and model health.
• Collaborate with ML engineers to accelerate model iteration while ensuring production safety, scalability, and cost efficiency.
• Enhance the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
• Spearhead architectural enhancements to create a more robust, user-friendly, scalable, and cost-effective online ML platform.
• Proven experience in building and managing production-quality online ML inference systems.
• Familiarity with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or comparable systems.
• Expertise in optimizing inference workloads through methods like dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
• Strong background in distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
• Proficient programming skills in Python, with hands-on experience working on production ML systems and high-scale services.
• Experience with PyTorch and contemporary model deployment workflows, including model packaging, validation, and lifecycle management for serving.
• Knowledge in designing infrastructure for safe model rollouts, canary testing, A/B experimentation, and automated rollbacks.
• Strong systems thinking and the ability to evaluate latency, throughput, reliability, scalability, and cost trade-offs.
• Demonstrated capability to guide technical direction and influence architectural decisions across teams without formal authority.
• Adequate knowledge of English for professional verbal and written communication.
• Equity awards.
• Participation in company incentive programs, such as annual discretionary bonuses or sales commissions.
• Comprehensive health, life, and disability insurance.
• Commute subsidy.
• Employee stock ownership.
• Competitive retirement and pension plans.
• Generous vacation and personal days.
• Leave and family care programs for new parents.
• Office snacks.
• Mental health and wellbeing programs and support.
• Employee Resource Groups.
• Global Employee Assistance Program.
• Training and development opportunities.
• Volunteering and donation matching programs.
Quora
Anthology Careers
AlertMedia
Vantor
Get handpicked remote jobs straight to your inbox weekly.