Senior Machine Learning Operations Engineer

Posted Jul 11

This is a fully remote position, open to applicants in Nebraska, +4 more states.

πŸ“‹ Description

β€’ Establish and manage BetMGM's ML platform on AWS (including SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (utilizing Snowpark ML, Cortex), supported by Terraform-managed infrastructure.

β€’ Create self-service frameworks that empower data scientists to deploy a model end-to-end without relying on a ticketing system β€” providing standardized project templates equipped with CI, drift monitoring, alerting, IaC, and Snowflake connectivity.

β€’ Design and manage batch scoring pipelines β€” utilizing SageMaker Batch Transform, orchestrated scoring via dbt against Snowflake, and Snowpark ML β€” with defined freshness and cost SLAs.

β€’ Design and oversee real-time inference pathways β€” leveraging SageMaker real-time endpoints, Lambda + Bedrock for GenAI, and API Gateway β€” adhering to specified latency budgets (generally under 100ms) and ensuring graceful degradation during high demand.

β€’ Take ownership of the feature store (using SageMaker Feature Store, Tecton, or Feast) ensuring online/offline parity β€” any training-serving skew is treated as an incident, not a compromise.

β€’ Develop CI/CD processes for ML β€” encompassing model registry, automated retraining triggers, model versioning, and tracking lineage from feature to training run to deployed model to live prediction.

β€’ Implement champion/challenger, shadow deployments, and canary releases as foundational platform features to prevent individual model teams from recreating these processes for each project.

β€’ Establish drift detection, data quality, and model performance monitoring (selecting from Evidently, Arize, or SageMaker Model Monitor β€” standardizing on one) with alerting routed to personnel capable of resolution.

β€’ Manage MLOps incident response β€” production model failures are classified as SEV events that require postmortem analysis.

β€’ Optimize endpoint configurations, batch caching, request batching, and autoscaling. Clearly define cost-per-prediction targets from the outset and ensure they are met.

β€’ Integrate LLM APIs (such as Bedrock, Anthropic, OpenAI) into production workflows β€” including RAG pipelines, agent evaluation frameworks, prompt versioning, and monitoring of costs and latency.


⛳️ Requirements

β€’ A BS or MS in Computer Science, Mathematics, Statistics, Machine Learning, or another STEM field β€” or equivalent real-world experience.

β€’ Over 5 years of experience delivering software in production environments β€” proficiency in Python, Docker, Kubernetes or ECS, CI/CD, and debugging distributed systems β€” including on-call responsibilities.

β€’ More than 3 years of experience managing ML in production β€” having owned a model that served real user traffic, with defined latency and cost parameters, and a runbook you authored.

β€’ In-depth experience with AWS services, particularly across the SageMaker ecosystem (Training, Endpoints, Batch Transform, Model Registry, Pipelines).

β€’ Proficient in Snowflake β€” including Snowpark ML, Cortex, dbt-orchestrated batch scoring, and RBAC for ML tasks.

β€’ Experience with IaC for ML β€” utilizing Terraform and SageMaker Pipelines or a comparable solution. No manual console deployments to production.

β€’ Background in feature store management β€” such as SageMaker Feature Store, Tecton, or Feast β€” with clear accountability for maintaining online/offline parity.

β€’ Familiarity with champion/challenger, shadow, and canary deployment strategies as practical experience, not just theoretical knowledge.

β€’ Experience with drift and model monitoring tools β€” such as Evidently, Arize, WhyLabs, or SageMaker Model Monitor β€” integrated into a notification system.

β€’ A software-engineering-first approach β€” treating ML systems as comprehensive systems rather than merely notebooks.


🏝️ Benefits

β€’ Medical, Dental, Vision, Life, and Disability Insurance.

β€’ 401(k) plan with company matching contributions.

β€’ Pre-tax spending accounts, including health care FSA and commuter savings.

β€’ Flexible paid time off policy.

β€’ Reimbursement for professional development and ongoing training opportunities.

β€’ Access to employee resource groups.

β€’ Swag, ticket giveaways, and additional perks!

People also viewed

24-MAG6 hours ago

ML Engineer

US flagNew York OnlyPart-timeMachine Learning Engineer$100 – $150/hour
ApplyView job
Aftershoot11 hours ago

Senior Machine Learning Engineer

IN flagIndia OnlyFull-timeMachine Learning Engineer
ApplyView job
Clinical Outcomes Solutions12 hours ago

Senior AI/ML Engineer

US flagUnited States, +1 more countryFull-timeMachine Learning Engineer
ApplyView job
Jalasoft1 day ago

Applied ML Engineer – Content Developer, Optimization & Foundations

MX flagMexico, +3 more countriesFreelanceMachine Learning Engineer
ApplyView job
Intetics1 day ago

Senior ML Engineer – 3 Month Project

DE flagGermany OnlyFreelanceMachine Learning Engineer
ApplyView job
Sistema Fibra1 day ago

PhD Scholarship Holder – Statistical Modeling, Causal Inference, and Machine Learning

BR flagBrazil OnlyFull-timeMachine Learning EngineerR$11k/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers