
Senior MLOps Engineer
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in California.
• Assist in transforming the models created by our ML Scientists, Data Scientists, and Perception Engineers into dependable, production-ready services.
• Engage with the infrastructure, pipelines, and tools that transition a model or an LLM/agent-driven workflow from a research notebook to a fully monitored deployment operating across various industry sectors.
• Sustain and enhance our model registry while constructing and troubleshooting deployment pipelines and cloud infrastructure, along with establishing model and pipeline monitoring and testing.
• Diagnose problems such as deployment failures, permission issues, or inconsistent environments.
• Contribute to wider automation efforts, facilitating deployment visibility and pipeline reliability that enable R&D, Software, Product, and Ops teams to collaborate effectively.
• Act as a primary communicator to ensure that R&D objectives and challenges are clearly understood by Software Engineering and DevOps teams.
• Collaborate with Scientists, working directly and iteratively with ML Scientists, Data Scientists, and Perception Engineers to convert experimental, research-focused code into reliable, scalable production services without hindering their research pace.
• A Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or a related discipline; typically requires 4+ years of professional experience in MLOps, ML platform engineering, or infrastructure engineering supporting machine learning teams.
• Strong practical knowledge of MLOps practices, capable of working autonomously across diverse production scenarios and escalating only truly complex or unclear issues.
• Proven experience collaborating directly with researchers or ML scientists, understanding research workflows, and translating them into reliable services and production models without creating bottlenecks.
• Proficient in Python with solid software engineering principles (testing, code review, version control).
• Practical experience with a leading cloud platform (e.g., AWS), infrastructure-as-code (Terraform), CI/CD tools (Github Actions), and containerization/orchestration technologies (e.g., Docker, Kubernetes).
• Experience in constructing and managing production ML pipelines and model registries, including model versioning and safer release procedures (e.g., canary deployments, rollbacks) across environments, as well as coordinating moderately complex, cross-functional infrastructure or deployment projects.
• Familiarity with building feedback loops from production back into training data, capturing human corrections as labels, and transforming retraining into a repeatable process.
• Understanding of experiment tracking, dataset/model versioning, and model documentation practices that support reproducible, auditable ML workflows is advantageous.
• Knowledge of computer vision or geospatial ML pipelines is a plus.
• Experience in deploying LLM/Agentic systems in production, including evaluation harnesses, prompt/tool/retrieval versioning, tracing, and token cost optimization is a plus.
• Skilled in creating data pipelines for relational databases (e.g., PostgreSQL) and API/GraphQL data layers (e.g., Hasura), and integrating external/third-party APIs into production workflows.
• A selection of various medical insurance plans, including options with an HSA and full premium coverage for yourself and your dependents.
• Complete dental and vision insurance coverage at no cost.
• Unlimited PTO: We genuinely prioritize work-life balance and mental well-being.
• Opportunities for autonomy and career advancement.
• A diverse, equitable, and inclusive culture: a workplace where your voice is valued.
Docket
Oxford Instruments plc
VIAFLOW®
Get handpicked remote jobs straight to your inbox weekly.