
Staff Machine Learning Engineer – Platform, MLOPS
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in Michigan.
• Take ownership of and manage the platform that underpins all ML and AI models at Credit Acceptance.
• Oversee the comprehensive deployment process for ML and GenAI models, encompassing training and inference pipelines, model registry and versioning, serving endpoints, and controlled promotion across development, QA, and production environments.
• Ensure runtime health for production models through consistent monitoring, alerting, drift and quality-regression detection, as well as managing latency and throughput objectives, capacity, autoscaling, and incident response via root cause analysis and corrective measures.
• Manage the agent runtime layer through the enterprise AI Gateway and MCP Gateway; transition agents onto the governed path while maintaining scoped, versioned, least-privilege tool surfaces.
• Supervise the production evaluation of agents, including online scoring, behavioral monitoring, drift detection, sampling and judge pipelines, as well as release gates.
• Develop and sustain observability and evaluation infrastructure, which includes multi-turn and multi-step agent tracing and telemetry capture, logging standards, pipeline configurations, and data contracts.
• Evaluate and manage platform unit economics, focusing on cost per inference, documentation, and interaction, while providing cost analysis for build-versus-buy and hosting decisions.
• Create reusable pipeline templates, deployment patterns, reference implementations, and internal tools to facilitate the adoption of the standard ML platform pathway.
• Collaborate with Cloud Engineering, Data Engineering, Security, and SRE on enterprise governance, identity, and observability strategies.
• Address AI-specific production incidents like prompt injection, rogue-agent cost spikes, data-classification exposure, delegation abuse, and model endpoint failures; ensure corrective actions are effectively implemented.
• Maintain architecture documentation and system diagrams related to the ML platform.
• Mentor engineers and interns on production ML practices and enhance operating standards through design and code reviews.
• Work remotely with occasional planned travel to the assigned office in Southfield, Michigan; may work at that office if requested by team members.
• Comply with company policies, processes, and legal regulations.
• Perform additional duties as assigned; attendance as needed by the department.
• Bachelor’s degree in Computer Science, Engineering, Statistics, or a relevant technical field with a minimum of 7 years of related experience, or a Master’s degree in one of those fields with at least 5 years of relevant experience.
• Over 5 years of experience in building and operating production ML or AI systems, with direct responsibility for at least two of the following: training or inference pipelines, model serving infrastructure, model registry and versioning, or production monitoring and alerting.
• Proven ownership of a production ML or AI service throughout its entire operational lifecycle.
• Proficient in Python and SQL, employing production-quality engineering practices such as version control, testing, code reviews, and CI/CD specifically applied to ML workloads.
• Hands-on experience with a cloud ML platform in production; AWS and Databricks are strongly preferred, encompassing model serving, job orchestration, and a model registry or experiment tracking system like MLflow.
• Understanding of the operational differences between LLM/GenAI workloads and traditional ML, including token costs, latency behavior, caching and batching, as well as non-deterministic outputs.
• Experience in running LLM or agent applications behind a gateway or proxy layer, including model routing and fallback, credential and key management, rate limiting, and budget enforcement.
• Familiarity with tool-calling architecture for agents, including the Model Context Protocol, MCP servers, tool scoping and authorization, and gateway-brokered tool access.
• Experience with containerization and infrastructure as code.
• Strong ability to communicate technical and non-technical trade-offs effectively in writing.
• Preferred: production experience with GPU-backed model serving, OpenTelemetry, and enterprise observability platforms such as Dynatrace; familiarity with model serving efficiency techniques, Databricks Unity Catalog, regulated-industry AI systems, agentic or multi-step AI systems, managed agent/tool gateways such as AWS Bedrock AgentCore Gateway, governed MCP servers, and emerging agent interoperability and identity standards like Agent2Agent (A2A).
• Ability to articulate complex technical information both verbally and in writing to various audiences, including senior leadership.
• Capacity to resolve problems at the source with effective, straightforward solutions.
• Prompt and efficient resolution of incidents, tasks, and projects.
• Demonstrated ability and enthusiasm for teaching others.
• Proficient in establishing relationships across different levels within the organization.
• Ability to prioritize and execute tasks in a high-pressure environment.
• Required degrees must be obtained from accredited institutions of higher education recognized by the Council for Higher Education Accreditation or equivalent.
• Annual variable bonus of cash and equity, ranging from 10-20%.
• 401(K) matching.
• Adoption assistance.
• Parental leave.
• Tuition reimbursement.
• Comprehensive medical, dental, and vision coverage.
• A variety of nonstandard benefits.
• Potential premium on top of the posted range for candidates located in San Francisco, Seattle, Boston, New York City, Los Angeles, and San Diego zones.
Planet Technologies
Northrop Grumman
Planet Technologies
M&T Bank
Get handpicked remote jobs straight to your inbox weekly.