
Senior MLOps, ML Platform Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in Mali, +1 more country.
• Develop and sustain ML training orchestration pipelines scheduled on an hourly, daily, and weekly basis.
• Implement mechanisms for retries, backfills, and idempotent execution.
• Design and manage model registry workflows, including versioning, lineage tracking, evaluation gates, and promotion procedures.
• Create isolated model environments for individual advertisers, ensuring namespace and configuration separation.
• Construct scalable refresh pipelines and publishing workflows tailored for serving infrastructure.
• Execute shadow mode and champion/challenger deployment strategies.
• Develop monitoring and alert systems for ML-specific metrics such as feature drift, prediction drift, training/serving skew, and calibration decay.
• Guarantee the reproducibility of ML workflows through containerized environments, pinned dependencies, and data snapshots.
• Oversee training and scoring costs across different tenants.
• Partner with DevOps and SRE engineers to enhance CI/CD and infrastructure automation.
• Create operational documentation and materials for platform handover.
• A minimum of 5 years of experience in MLOps, ML platform engineering, or infrastructure engineering that supports production ML systems.
• Proficient in Python with a background in building platform-level tools and automation.
• Practical experience with Kubernetes and Docker.
• Familiarity in constructing CI/CD pipelines for ML workloads.
• Hands-on experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or comparable orchestration and ML lifecycle platforms.
• Experience with ML platforms and model lifecycle tools like Vertex AI, MLflow, or Kubeflow.
• Strong grasp of ML observability, including drift detection, monitoring train/serve skew, and incident response.
• Background in designing or supporting multi-tenant ML systems and isolated model environments.
• Experience with cloud platforms, ideally GCP.
• Knowledge of infrastructure-as-code tools such as Terraform.
• Proficiency in Linux environments.
• Understanding of the ML lifecycle and productionization processes.
• Upper-Intermediate English proficiency or better.
• Opportunity for remote work.
• Chance to engage in innovative ML infrastructure projects.
• Collaboration with seasoned engineers.
• Ability to influence architectural decisions.
• Long-term strategic engagement.
The College Board
KnowBe4
SeatGeek
AIS (Applied Information Sciences)
Get handpicked remote jobs straight to your inbox weekly.