
Manager, Machine Learning – AI Modeling and Operation
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the ML model lifecycle, encompassing training pipelines, model registry, deployment, monitoring, and implementing guardrails.
• Develop and sustain CI/CD for ML, which includes automated testing, evaluation, and model promotion across different environments.
• Enhance observability across AI services by tracking latency, detecting drift, monitoring costs, and setting up alerts.
• Define SLOs/SLIs for AI services and spearhead incident response for ML-related production challenges.
• Simplify processes through automation, thoughtful system design, and reducing complexity.
• Lead and expand an existing team of machine learning engineers.
• Offer hands-on coaching, performance assessments, and opportunities for growth.
• Cultivate a collaborative, inclusive, and high-ownership team culture.
• Collaborate with engineering squads and Product teams to ensure AI services are ready for production, operationally sound, and easily observable.
• Convey complex technical issues effectively to both technical and non-technical audiences.
• Manage integrations with leading model providers, addressing model routing, load balancing, and fallback strategies.
• Drive the development of AI analytics dashboards for platform health and usage insights.
• Enhance latency, service availability, developer experience, and integration usability.
• Guide architectural decisions to ensure platform scalability, reliability, and alignment with Workiva’s technical vision.
• Bachelor's degree in Computer Science, Engineering, Data Science, or a relevant combination of education and experience.
• Over 7 years of overall experience in software engineering and/or Machine Learning.
• Minimum of 2 years of focused experience as an Engineering Manager.
• Familiarity with ML pipeline orchestration tools like ClearML, Kubeflow, Airflow, or similar platforms.
• Practical experience with Kubernetes, microservices, container orchestration, and infrastructure-as-code.
• Proven track record of enhancing reliability and availability metrics for production ML systems.
• Ability to manage senior individual contributors, resolve technical disputes, and foster psychological safety and high performance.
• Strong leadership capabilities within an Agile/Sprint working environment.
• Experience managing production ML systems in cloud environments such as AWS, Azure, or GCP.
• Willingness to travel up to 15% for team and corporate meetings.
• Dependable internet access for remote working opportunities.
• Preferred: Master’s degree in Computer Science, Engineering, Data Science, or equivalent experience.
• Preferred: Experience with Generative AI concepts, including RAG and Agentic frameworks.
• Preferred: Background in building model evaluation or quality measurement systems.
• Preferred: Knowledge of cost optimization for GPU/model serving workloads.
• Preferred: Familiarity with observability tools such as Datadog, Prometheus, or Grafana.
• A discretionary bonus typically awarded annually.
• Restricted Stock Units granted upon hiring.
• 401(k) matching program.
• Comprehensive employee benefits package.
• Opportunity to work remotely from anywhere within the country of employment.
• Travel up to 15% for team and corporate meetings.
Torc Robotics
StackAdapt
Torc Robotics
Get handpicked remote jobs straight to your inbox weekly.