
Senior ML / AI Engineer
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Design and implement agentic systems utilizing dynamic tool-calling agents, structured Pydantic outputs, managed tool catalogs, per-step verification, grounding checks, LLM-as-judge, human-in-the-loop, and bounded auditable loops.
• Develop the RAG layer, which encompasses document ingestion, chunking, embeddings, hybrid retrieval, reranking, access scoping, and injection/poisoning defenses.
• Manage self-hosted inference at scale with vLLM and embedding/rerank services on EKS GPU nodes.
• Enhance inference throughput through continuous batching, prefix caching, quantization, and KV-cache tuning.
• Implement fair-share concurrency across tenants.
• Oversee the MLOps and governance plane, which includes offline evaluation harnesses, golden sets, quality scoring, canary/gray releases, auto-rollback, cost/budget governance, and observability.
• Ensure production safety of systems through durable state, circuit breakers, retries, PHI-safe logging/redaction, RBAC, fail-closed defaults, and horizontal AWS scalability via Terraform.
• Expand the platform into new domains from problem framing to governed, evaluated, and deployed services.
• Guide engineers in developing initiatives based on the shared platform foundation.
• Maintain clean, typed, and tested code while actively participating in thoughtful design reviews.
• Balance autonomy and determinism for high-stakes tasks.
• Over 5 years of experience in building and deploying ML/AI systems in production, with recent hands-on work in LLM applications over the past 1–2 years.
• Proficient in Python (typed, tested, production-level) and possess strong software engineering fundamentals.
• Comfortable working across an asynchronous web service, data layer, and infrastructure.
• Practical expertise in prompting, structured/function-calling outputs, RAG, embeddings, vector search, retrieval quality, agent/tool-use loops, grounding, hallucination control, evaluation sets, and guardrails.
• Cloud and MLOps experience on AWS or a similar platform, including containers, Kubernetes, Terraform, CI/CD, observability, and model-serving cost/performance optimization.
• Proven track record in ensuring reliability, encompassing state durability, failure management, scaling, and debugging production incidents.
• Strong written communication skills and sound judgment in ambiguous situations.
• Experience in self-hosting/optimizing open-weight models such as vLLM or TGI on GPUs.
• Familiarity with embeddings/rerank serving, such as bge or Infinity.
• Experience with LangGraph, LangChain, or similar agent frameworks and multi-agent orchestration.
• Knowledge of regulated data practices involving HIPAA, SOC 2, ISO 27001, PHI/PII handling, RBAC, and audit trails.
• Proficiency in pgvector/Postgres, MongoDB/DocumentDB, SQS, and Cognito/OIDC.
• Experience with OpenTelemetry, Prometheus/Grafana, and frontend development using React/TypeScript.
• Familiarity with Amazon Bedrock or a multi-provider abstraction.
• Experience with on-premises or air-gapped deployment.
• Background in healthcare, insurance, or claims domains, including familiarity with X12 835, EOB, or benefits, is highly valued.
• Must be authorized to work for any employer in the United States; this position does not provide visa sponsorship.
• Comprehensive coverage of medical, dental, and vision benefits for both employees and their dependents.
• 401k match up to 4%.
Alzheimer's Association®
Capital One
Capital One
Get handpicked remote jobs straight to your inbox weekly.