
MLOps, ML Platform Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Design and manage ML infrastructure for data processing, training, serving, and inference systems.
• Implement reproducible training and evaluation pipelines that include versioning, scheduling, and artifact tracking.
• Optimize GPU and CPU workloads, oversee clusters, and enhance compute costs through rightsizing, spot scheduling, and caching strategies.
• Operate production APIs facilitating low-latency inference with features like autoscaling, blue-green or canary rollouts, and rollback safety.
• Define and take ownership of Service Level Objectives (SLOs); instrument pipelines and services to monitor latency, costs, drift, and data quality.
• Manage Identity and Access Management (IAM), secrets, and container security.
• Automate deployment pipelines utilizing CI/CD and infrastructure as code methodologies.
• Collaborate with research scientists and AI engineers to transition models from experimentation to production.
• Create templates, runbooks, and internal tools to ensure ML workflows are repeatable, secure, and efficient.
• Minimum of 4 years of experience in ML platform, DevOps, or infrastructure engineering.
• In-depth knowledge of Kubernetes, CI/CD, containerization, and cloud infrastructure (AWS, GCP, or Azure).
• Practical experience in managing GPU clusters and training/inference pipelines.
• Familiarity with data orchestration and storage formats such as Delta, Parquet, Polars, and Spark.
• Demonstrated ability to deploy and maintain production ML systems adhering to SLOs.
• Strong Python programming skills and proficiency in infrastructure as code and automation practices.
• Experience with observability and cost optimization at scale.
• Background in real-time or low-latency model serving (REST, gRPC).
• Exposure to model registry and promotion workflows.
• Understanding of data quality, lineage, and curation pipelines.
• Experience in sports analytics or similar high-volume data fields.
• Familiarity with integrating LLM workflows or evaluation pipelines.
• Competitive Salary and Bonus Plan.
• Comprehensive health insurance plan.
• Retirement savings plan (401k) with company match.
• Remote working environment.
• A flexible, unlimited time off policy.
• Generous paid holiday schedule - 13 in total including the Monday after the Super Bowl.
• Annual performance bonus.
• Benefits and/or other applicable incentive compensation plans.
EverCommerce
ARETUM
Civica US
Vannevar Labs
Get handpicked remote jobs straight to your inbox weekly.