Staff ML Engineer

Posted 3 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee the design, delivery, and functioning of production-grade AI systems.

• Develop services and high-throughput pipelines for a text analytics platform and AI products.

• Take ownership of technical delivery from prototype development to deployment and ongoing production support.

• Collaborate with AI Scientists and Product teams to establish requirements, plan implementations, and address cross-team dependencies.

• Assess AI solutions for their suitability in production and identify associated technical risks.

• Select architectures that fulfill quality, reliability, and cost requirements.

• Design and construct high-volume AI services and processing pipelines with features such as fault tolerance, backpressure, retries, idempotency, and partial-failure recovery.

• Spearhead the advancement of Python services and a Databricks-based platform for distributed processing, model serving, and ML/LLM integration.

• Set standards for automated testing, CI/CD, model and prompt versioning, load testing, controlled rollouts, and rollback procedures.

• Develop capabilities for evaluating and monitoring AI quality regressions, service reliability, throughput, latency, and inference costs.

• Collaborate with Product and Responsible AI teams regarding release criteria, model validation, data privacy, security, and governance.

• Optimize processing and inference workloads for quality, throughput, latency, capacity, and cost efficiency.

• Guide and mentor engineers while leading architecture and code reviews.


⛳️ Requirements

• Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.

• 8+ years of professional software engineering experience.

• A minimum of 3 years of experience managing ML or LLM systems in production along with their operational support.

• Demonstrated capability to independently lead intricate technical initiatives from requirements gathering to production rollout.

• Advanced proficiency in Python for production services and data processing tasks.

• Strong SQL capabilities.

• Experience in designing and managing high-throughput distributed systems.

• Comprehensive understanding of failure recovery, multi-tenancy, and capacity planning.

• Practical experience in deploying and operating LLM-based applications, including evaluation, output validation, observability, and cost management.

• Strong adherence to production engineering best practices, including automated testing, CI/CD, monitoring, incident response, and root-cause analysis.

• Evidence of technical leadership through system design, hands-on implementation, code reviews, and mentorship activities.

• Ability to clearly communicate technical decisions and tradeoffs to engineering, research, product, and governance stakeholders.

• Preferred: Familiarity with Databricks or similar cloud-based data and AI platforms.

• Preferred: Experience with NLP, text analytics, or large-scale processing of unstructured data.

• Preferred: Background in building shared infrastructure for inference, evaluation, and model lifecycle management.

• Preferred: Knowledge of retrieval-augmented generation, semantic search, and LLM orchestration frameworks.

• Preferred: Experience with speech-to-text, speaker diarization, or conversational audio processing.

• Preferred: Experience in deploying and operating cloud-native services on AWS or Azure.

• Preferred: Experience in healthcare or other regulated environments, including management of sensitive data, auditability, and model governance.


🏝️ Benefits

• Competitive benefits package.

• Discretionary bonus or commission based on successful outcomes.

• Reasonable accommodations for qualified individuals with disabilities or disabled veterans throughout the hiring process.

People also viewed

Shield AI21 hours ago

Staff Deep Learning Engineer, State Estimation

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $300k/year
ApplyView job
Weekday (YC W21)1 day ago

ML Engineer

IN flagIndia OnlyFull-timeMachine Learning Engineer₹2.5M – ₹5M/year
ApplyView job
Roadpass Digital1 day ago

Senior AI/ML Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer
ApplyView job
MWDN1 day ago

AI/ML Engineer

HR flagCroatia OnlyFull-timeMachine Learning Engineer
ApplyView job
Quora1 day ago

Software Engineer, New Grad – Machine Learning Platform

US flagUnited States, +1 more countryFull-timeMachine Learning Engineer$97.6k – $139k/year
ApplyView job
Amgen1 day ago

Principal Machine Learning Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer$187.4k – $253.5k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers