
Staff ML Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Oversee the design, delivery, and functioning of production-grade AI systems.
• Develop services and high-throughput pipelines for a text analytics platform and AI products.
• Take ownership of technical delivery from prototype development to deployment and ongoing production support.
• Collaborate with AI Scientists and Product teams to establish requirements, plan implementations, and address cross-team dependencies.
• Assess AI solutions for their suitability in production and identify associated technical risks.
• Select architectures that fulfill quality, reliability, and cost requirements.
• Design and construct high-volume AI services and processing pipelines with features such as fault tolerance, backpressure, retries, idempotency, and partial-failure recovery.
• Spearhead the advancement of Python services and a Databricks-based platform for distributed processing, model serving, and ML/LLM integration.
• Set standards for automated testing, CI/CD, model and prompt versioning, load testing, controlled rollouts, and rollback procedures.
• Develop capabilities for evaluating and monitoring AI quality regressions, service reliability, throughput, latency, and inference costs.
• Collaborate with Product and Responsible AI teams regarding release criteria, model validation, data privacy, security, and governance.
• Optimize processing and inference workloads for quality, throughput, latency, capacity, and cost efficiency.
• Guide and mentor engineers while leading architecture and code reviews.
• Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
• 8+ years of professional software engineering experience.
• A minimum of 3 years of experience managing ML or LLM systems in production along with their operational support.
• Demonstrated capability to independently lead intricate technical initiatives from requirements gathering to production rollout.
• Advanced proficiency in Python for production services and data processing tasks.
• Strong SQL capabilities.
• Experience in designing and managing high-throughput distributed systems.
• Comprehensive understanding of failure recovery, multi-tenancy, and capacity planning.
• Practical experience in deploying and operating LLM-based applications, including evaluation, output validation, observability, and cost management.
• Strong adherence to production engineering best practices, including automated testing, CI/CD, monitoring, incident response, and root-cause analysis.
• Evidence of technical leadership through system design, hands-on implementation, code reviews, and mentorship activities.
• Ability to clearly communicate technical decisions and tradeoffs to engineering, research, product, and governance stakeholders.
• Preferred: Familiarity with Databricks or similar cloud-based data and AI platforms.
• Preferred: Experience with NLP, text analytics, or large-scale processing of unstructured data.
• Preferred: Background in building shared infrastructure for inference, evaluation, and model lifecycle management.
• Preferred: Knowledge of retrieval-augmented generation, semantic search, and LLM orchestration frameworks.
• Preferred: Experience with speech-to-text, speaker diarization, or conversational audio processing.
• Preferred: Experience in deploying and operating cloud-native services on AWS or Azure.
• Preferred: Experience in healthcare or other regulated environments, including management of sensitive data, auditability, and model governance.
• Competitive benefits package.
• Discretionary bonus or commission based on successful outcomes.
• Reasonable accommodations for qualified individuals with disabilities or disabled veterans throughout the hiring process.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.