
Senior MLOps Engineer – US East Coast Time Zone
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Alaska, +7 more states.
• Oversee the development and enhancement of the environment, platform capabilities, and operational foundations that support machine learning workflows throughout Oura.
• Lead the design and integration of workflows and tools for dependable ML training, orchestration, and deployment.
• Collaborate with data scientists and engineers to refine the complete ML lifecycle from experimentation and training to deployment and governance.
• Establish and advance model governance practices, which include reproducibility, lineage, access controls, and operational standards.
• Standardize ML tools and workflows such as experiment tracking, model packaging, and promotion procedures.
• Work together across business domains to onboard use cases and prioritize collective ML platform enhancements.
• Address reliability, scalability, and cost-efficiency challenges within cloud ML infrastructure.
• Promote standards, automation, infrastructure-as-code, CI/CD, and documentation among data science teams.
• Enhance observability and automation of infrastructure for ML workflows.
• Contribute to workflow orchestration and platform integrations for model training and batch inference.
• Align ML systems with wider data platform and governance practices.
• Develop best practices for creating, deploying, and maintaining production ML systems.
• A minimum of 5 years of experience in MLOps, machine learning engineering, platform engineering, data engineering, or a closely related discipline.
• Practical experience managing production workloads in AWS.
• Solid understanding of cloud infrastructure concepts.
• Comprehensive knowledge of the machine learning lifecycle, encompassing training workflows, deployment patterns, monitoring, and model maintenance.
• Acquainted with data science workflows and experimentation/model-operations tools like MLflow.
• Familiar with workflow orchestration, infrastructure-as-code, and CI/CD practices tailored for ML or data platforms.
• Knowledge of secure access patterns, governance controls, and shared cloud or data platform services.
• Experience in building or supporting production-grade ML workflows centered on reliability, reproducibility, and maintainability.
• Capability to drive standards and enhancements across various teams and business domains.
• Excellent communication and collaboration abilities with both technical and non-technical stakeholders.
• Ability to function within a distributed team while demonstrating ownership and autonomy.
• Competitive salary and equity packages.
• Health, dental, and vision insurance, along with mental health resources.
• An Oura Ring for yourself plus employee discounts for friends and family.
• 20 days of paid time off, plus 13 paid holidays and 8 days of flexible wellness time off.
• Paid sick leave and parental leave.
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.