Remotery

Lead Software Platform Engineer, MLOps

Posted Aug 4

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Take full responsibility for the technical architecture of the AI/ML platform along with its customer-facing service and API components.

• Manage the comprehensive model and prompt lifecycle across Databricks MLflow and AWS Bedrock, encompassing registration, versioning, staged promotion, rollback, and multi-model serving.

• Develop inference infrastructure for both real-time and batch AI workloads, addressing aspects such as routing, batching, caching, concurrency control, GPU capacity planning, large binary inputs, and graceful degradation.

• Seamlessly integrate AI models and LLMs into production systems utilizing RAG, tool and function calling, MCP-based tooling, and agent runtimes.

• Architect platform security measures, including guardrails, defenses against prompt injection and tool abuse, handling of PII and PHI, and maintaining tenant data boundaries.

• Create AI evaluation and quality infrastructure, featuring offline and online evaluation harnesses, golden datasets, CI regression gates, A/B and shadow deployment, along with drift and hallucination detection.

• Implement monitoring, alerting, logging, distributed tracing, and SLI/SLO/SLA practices for the AI platform.

• Design systems for reproducibility and lineage within validated pharma environments, including versioned data, code, prompts, model artifacts, and audit trails.

• Contribute to infrastructure-as-code and deployment automation through CloudFormation and AWS CDK.

• Ensure production readiness in collaboration with Applied AI, data engineering, and platform teams, focusing on performance, reliability, cost efficiency, incident response, and runbooks.

• Serve as a Subject Matter Expert (SME) and design authority, leading design reviews, drafting reference architectures and technical documentation, and mentoring engineers.

• Assess emerging AI infrastructure frameworks, serving runtimes, model providers, and data types, making informed build-versus-buy decisions.


⛳️ Requirements

• Over 10 years of professional experience in software and infrastructure engineering, focusing on designing, building, and scaling distributed cloud-native production systems.

• Demonstrated technical leadership or architectural experience with accountability for system design, scalability, performance, and cost optimization.

• Experience in embedding security within multi-tenant platforms, including tenant authorization boundaries and handling of PII/PHI.

• Knowledge of LLM-specific risks, such as prompt injection and tool abuse.

• Extensive production experience in developing AI/ML infrastructure as a multi-tenant product for external users.

• In-depth hands-on experience in deploying LLM-based systems in production, including RAG, retrieval and embedding design, prompt and model versioning, and tool or function calling.

• Proficient in TypeScript and Python coding for creating robust APIs and backend services.

• Production experience with model registries and serving stacks, preferably with Databricks MLflow.

• Familiarity with AI evaluation release gates, regression gates, and monitoring for drift or quality.

• Skilled in API-first design, REST, and OpenAPI.

• Solid understanding of AWS and Docker.

• Experience with CloudFormation or AWS CDK, CI/CD pipelines, and deployment automation.

• Competence in defining observability and SLI/SLO/SLA practices, including monitoring, alerting, and distributed tracing.

• Ability to effectively communicate with customers and cross-functional teams, influence technical direction, and provide mentorship to engineers.

• Nice to have: familiarity with emerging LLM frameworks, production agentic orchestration, MCP, LLM cost monitoring, multimodal inputs, fine-tuning or model optimization, experience in regulated or validated environments, and knowledge in scientific or life sciences domains.

• Please note that visa sponsorship is not currently available for this position.


🏝️ Benefits

• 100% employer-paid benefits for all eligible employees and their immediate family members.

• Unlimited paid time off (PTO).

• 401K plan.

• Flexible working arrangements allowing for remote work.

• Company-paid Life Insurance, LTD/STD.

• A culture of continuous improvement, fostering career growth and coaching opportunities.

People also viewed

Agility Technologies Inc11 hours ago

Databricks Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer
ApplyView job
American College of Education14 hours ago

Power Platform Developer

US flagUnited States OnlyFull-timePlatform Engineer$107k/year
ApplyView job
First Due16 hours ago

Director of Platform Engineering

US flagUnited States OnlyFull-timePlatform Engineer$240k/year
ApplyView job
Faire17 hours ago

Senior Staff Machine Learning Platform Engineer

US flagArizona, +30 more statesFull-timePlatform Engineer$295k – $405.5k/year
ApplyView job
Faire17 hours ago

Senior Staff Machine Learning Platform Engineer

CA flagCanada OnlyFull-timePlatform EngineerC$248k – C$341k/year
ApplyView job
Faire17 hours ago

Staff Machine Learning Platform Engineer

US flagArizona, +29 more statesFull-timePlatform Engineer$246.5k – $339k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers