
Senior MLOps, ML Platform Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Mali, +1 more country.
• Develop and sustain ML training orchestration pipelines operating on hourly, daily, and weekly schedules
• Implement mechanisms for retries, backfills, and idempotent execution
• Design and facilitate model registry workflows, including versioning, lineage tracking, evaluation gates, and promotion procedures
• Create isolated model environments for each advertiser, ensuring namespace and configuration separation
• Construct scalable refresh pipelines and publishing workflows for the serving infrastructure
• Execute shadow mode and champion/challenger deployment strategies
• Establish monitoring and alerting systems for ML-specific metrics, such as feature drift, prediction drift, training/serving skew, and calibration decay
• Guarantee reproducibility of ML workflows through the use of containerized environments, pinned dependencies, and data snapshots
• Track training and scoring expenditures across different tenants
• Collaborate with DevOps and SRE engineers to enhance CI/CD and automate infrastructure
• Create operational documentation and prepare platform handover materials
• Over 5 years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
• Proficient in Python with experience in building platform-level tools and automation
• Practical experience with Kubernetes and Docker
• Proven experience in developing CI/CD pipelines for ML workloads
• Hands-on experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms
• Familiarity with ML platforms and model lifecycle tools like Vertex AI, MLflow, or Kubeflow
• Strong grasp of ML observability concepts, including drift detection, monitoring train/serve skew, and incident response
• Experience in designing or supporting multi-tenant ML systems and isolated model environments
• Background in working with cloud platforms, preferably GCP
• Familiarity with infrastructure-as-code tools such as Terraform
• Experience in Linux environments
• Understanding of the ML lifecycle and productionization processes
• Upper-Intermediate English proficiency or higher
• Option for remote work
• Opportunity to engage in innovative ML infrastructure projects
• Collaboration with skilled engineers
• Ability to influence architectural decisions
• Long-term strategic engagement
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.