
Senior Machine Learning Operations Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in Nebraska, +4 more states.
• Establish and manage BetMGM's ML platform on AWS (including SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (utilizing Snowpark ML, Cortex), supported by Terraform-managed infrastructure.
• Create self-service frameworks that empower data scientists to deploy a model end-to-end without relying on a ticketing system — providing standardized project templates equipped with CI, drift monitoring, alerting, IaC, and Snowflake connectivity.
• Design and manage batch scoring pipelines — utilizing SageMaker Batch Transform, orchestrated scoring via dbt against Snowflake, and Snowpark ML — with defined freshness and cost SLAs.
• Design and oversee real-time inference pathways — leveraging SageMaker real-time endpoints, Lambda + Bedrock for GenAI, and API Gateway — adhering to specified latency budgets (generally under 100ms) and ensuring graceful degradation during high demand.
• Take ownership of the feature store (using SageMaker Feature Store, Tecton, or Feast) ensuring online/offline parity — any training-serving skew is treated as an incident, not a compromise.
• Develop CI/CD processes for ML — encompassing model registry, automated retraining triggers, model versioning, and tracking lineage from feature to training run to deployed model to live prediction.
• Implement champion/challenger, shadow deployments, and canary releases as foundational platform features to prevent individual model teams from recreating these processes for each project.
• Establish drift detection, data quality, and model performance monitoring (selecting from Evidently, Arize, or SageMaker Model Monitor — standardizing on one) with alerting routed to personnel capable of resolution.
• Manage MLOps incident response — production model failures are classified as SEV events that require postmortem analysis.
• Optimize endpoint configurations, batch caching, request batching, and autoscaling. Clearly define cost-per-prediction targets from the outset and ensure they are met.
• Integrate LLM APIs (such as Bedrock, Anthropic, OpenAI) into production workflows — including RAG pipelines, agent evaluation frameworks, prompt versioning, and monitoring of costs and latency.
• A BS or MS in Computer Science, Mathematics, Statistics, Machine Learning, or another STEM field — or equivalent real-world experience.
• Over 5 years of experience delivering software in production environments — proficiency in Python, Docker, Kubernetes or ECS, CI/CD, and debugging distributed systems — including on-call responsibilities.
• More than 3 years of experience managing ML in production — having owned a model that served real user traffic, with defined latency and cost parameters, and a runbook you authored.
• In-depth experience with AWS services, particularly across the SageMaker ecosystem (Training, Endpoints, Batch Transform, Model Registry, Pipelines).
• Proficient in Snowflake — including Snowpark ML, Cortex, dbt-orchestrated batch scoring, and RBAC for ML tasks.
• Experience with IaC for ML — utilizing Terraform and SageMaker Pipelines or a comparable solution. No manual console deployments to production.
• Background in feature store management — such as SageMaker Feature Store, Tecton, or Feast — with clear accountability for maintaining online/offline parity.
• Familiarity with champion/challenger, shadow, and canary deployment strategies as practical experience, not just theoretical knowledge.
• Experience with drift and model monitoring tools — such as Evidently, Arize, WhyLabs, or SageMaker Model Monitor — integrated into a notification system.
• A software-engineering-first approach — treating ML systems as comprehensive systems rather than merely notebooks.
• Medical, Dental, Vision, Life, and Disability Insurance.
• 401(k) plan with company matching contributions.
• Pre-tax spending accounts, including health care FSA and commuter savings.
• Flexible paid time off policy.
• Reimbursement for professional development and ongoing training opportunities.
• Access to employee resource groups.
• Swag, ticket giveaways, and additional perks!
Docket
BJAK
Oxford Instruments plc
VIAFLOW®
Get handpicked remote jobs straight to your inbox weekly.