
Lead Software Platform Engineer, MLOps
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in California, +1 more state.
• Take full responsibility for the technical architecture of the AI/ML platform along with its customer-facing service and API components.
• Manage the comprehensive model and prompt lifecycle across Databricks MLflow and AWS Bedrock, encompassing registration, versioning, staged promotion, rollback, and multi-model serving.
• Develop inference infrastructure for both real-time and batch AI workloads, addressing aspects such as routing, batching, caching, concurrency control, GPU capacity planning, large binary inputs, and graceful degradation.
• Seamlessly integrate AI models and LLMs into production systems utilizing RAG, tool and function calling, MCP-based tooling, and agent runtimes.
• Architect platform security measures, including guardrails, defenses against prompt injection and tool abuse, handling of PII and PHI, and maintaining tenant data boundaries.
• Create AI evaluation and quality infrastructure, featuring offline and online evaluation harnesses, golden datasets, CI regression gates, A/B and shadow deployment, along with drift and hallucination detection.
• Implement monitoring, alerting, logging, distributed tracing, and SLI/SLO/SLA practices for the AI platform.
• Design systems for reproducibility and lineage within validated pharma environments, including versioned data, code, prompts, model artifacts, and audit trails.
• Contribute to infrastructure-as-code and deployment automation through CloudFormation and AWS CDK.
• Ensure production readiness in collaboration with Applied AI, data engineering, and platform teams, focusing on performance, reliability, cost efficiency, incident response, and runbooks.
• Serve as a Subject Matter Expert (SME) and design authority, leading design reviews, drafting reference architectures and technical documentation, and mentoring engineers.
• Assess emerging AI infrastructure frameworks, serving runtimes, model providers, and data types, making informed build-versus-buy decisions.
• Over 10 years of professional experience in software and infrastructure engineering, focusing on designing, building, and scaling distributed cloud-native production systems.
• Demonstrated technical leadership or architectural experience with accountability for system design, scalability, performance, and cost optimization.
• Experience in embedding security within multi-tenant platforms, including tenant authorization boundaries and handling of PII/PHI.
• Knowledge of LLM-specific risks, such as prompt injection and tool abuse.
• Extensive production experience in developing AI/ML infrastructure as a multi-tenant product for external users.
• In-depth hands-on experience in deploying LLM-based systems in production, including RAG, retrieval and embedding design, prompt and model versioning, and tool or function calling.
• Proficient in TypeScript and Python coding for creating robust APIs and backend services.
• Production experience with model registries and serving stacks, preferably with Databricks MLflow.
• Familiarity with AI evaluation release gates, regression gates, and monitoring for drift or quality.
• Skilled in API-first design, REST, and OpenAPI.
• Solid understanding of AWS and Docker.
• Experience with CloudFormation or AWS CDK, CI/CD pipelines, and deployment automation.
• Competence in defining observability and SLI/SLO/SLA practices, including monitoring, alerting, and distributed tracing.
• Ability to effectively communicate with customers and cross-functional teams, influence technical direction, and provide mentorship to engineers.
• Nice to have: familiarity with emerging LLM frameworks, production agentic orchestration, MCP, LLM cost monitoring, multimodal inputs, fine-tuning or model optimization, experience in regulated or validated environments, and knowledge in scientific or life sciences domains.
• Please note that visa sponsorship is not currently available for this position.
• 100% employer-paid benefits for all eligible employees and their immediate family members.
• Unlimited paid time off (PTO).
• 401K plan.
• Flexible working arrangements allowing for remote work.
• Company-paid Life Insurance, LTD/STD.
• A culture of continuous improvement, fostering career growth and coaching opportunities.
Agility Technologies Inc
American College of Education
First Due
Faire
Get handpicked remote jobs straight to your inbox weekly.