Remotery

Senior Machine Learning Operations Engineer

Posted Jul 17

This is a fully remote position, open to applicants in Nebraska, +4 more states.

📋 Description

• Establish and manage BetMGM's ML platform on AWS (including SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (utilizing Snowpark ML, Cortex), supported by Terraform-managed infrastructure.

• Create self-service frameworks that empower data scientists to deploy a model end-to-end without relying on a ticketing system — providing standardized project templates equipped with CI, drift monitoring, alerting, IaC, and Snowflake connectivity.

• Design and manage batch scoring pipelines — utilizing SageMaker Batch Transform, orchestrated scoring via dbt against Snowflake, and Snowpark ML — with defined freshness and cost SLAs.

• Design and oversee real-time inference pathways — leveraging SageMaker real-time endpoints, Lambda + Bedrock for GenAI, and API Gateway — adhering to specified latency budgets (generally under 100ms) and ensuring graceful degradation during high demand.

• Take ownership of the feature store (using SageMaker Feature Store, Tecton, or Feast) ensuring online/offline parity — any training-serving skew is treated as an incident, not a compromise.

• Develop CI/CD processes for ML — encompassing model registry, automated retraining triggers, model versioning, and tracking lineage from feature to training run to deployed model to live prediction.

• Implement champion/challenger, shadow deployments, and canary releases as foundational platform features to prevent individual model teams from recreating these processes for each project.

• Establish drift detection, data quality, and model performance monitoring (selecting from Evidently, Arize, or SageMaker Model Monitor — standardizing on one) with alerting routed to personnel capable of resolution.

• Manage MLOps incident response — production model failures are classified as SEV events that require postmortem analysis.

• Optimize endpoint configurations, batch caching, request batching, and autoscaling. Clearly define cost-per-prediction targets from the outset and ensure they are met.

• Integrate LLM APIs (such as Bedrock, Anthropic, OpenAI) into production workflows — including RAG pipelines, agent evaluation frameworks, prompt versioning, and monitoring of costs and latency.


⛳️ Requirements

• A BS or MS in Computer Science, Mathematics, Statistics, Machine Learning, or another STEM field — or equivalent real-world experience.

• Over 5 years of experience delivering software in production environments — proficiency in Python, Docker, Kubernetes or ECS, CI/CD, and debugging distributed systems — including on-call responsibilities.

• More than 3 years of experience managing ML in production — having owned a model that served real user traffic, with defined latency and cost parameters, and a runbook you authored.

• In-depth experience with AWS services, particularly across the SageMaker ecosystem (Training, Endpoints, Batch Transform, Model Registry, Pipelines).

• Proficient in Snowflake — including Snowpark ML, Cortex, dbt-orchestrated batch scoring, and RBAC for ML tasks.

• Experience with IaC for ML — utilizing Terraform and SageMaker Pipelines or a comparable solution. No manual console deployments to production.

• Background in feature store management — such as SageMaker Feature Store, Tecton, or Feast — with clear accountability for maintaining online/offline parity.

• Familiarity with champion/challenger, shadow, and canary deployment strategies as practical experience, not just theoretical knowledge.

• Experience with drift and model monitoring tools — such as Evidently, Arize, WhyLabs, or SageMaker Model Monitor — integrated into a notification system.

• A software-engineering-first approach — treating ML systems as comprehensive systems rather than merely notebooks.


🏝️ Benefits

• Medical, Dental, Vision, Life, and Disability Insurance.

• 401(k) plan with company matching contributions.

• Pre-tax spending accounts, including health care FSA and commuter savings.

• Flexible paid time off policy.

• Reimbursement for professional development and ongoing training opportunities.

• Access to employee resource groups.

• Swag, ticket giveaways, and additional perks!

People also viewed

Docket1 day ago

Junior Machine Learning Engineer

Anywhere in the WorldFull-timeMachine Learning Engineer
ApplyView job
BJAK1 day ago

Senior Machine Learning Engineer

DE flagGermany OnlyFull-timeMachine Learning Engineer
ApplyView job
Oxford Instruments plc1 day ago

Senior Machine Learning Software Engineer

GB flagUnited Kingdom OnlyFull-timeMachine Learning Engineer
ApplyView job
VIAFLOW®1 day ago

Engenheira de Machine Learning Sênior

BR flagBrazil OnlyFull-timeMachine Learning Engineer
ApplyView job
Hire Hangar Global1 day ago

Machine Learning Engineer

CO flagColombia OnlyFreelanceMachine Learning Engineer$2,500 – $4,000/month
ApplyView job
Grailed1 day ago

Staff Machine Learning Engineer

US flagCalifornia, +1 more stateFull-timeMachine Learning Engineer$159k – $233.8k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers