
Staff ML Platform Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Establish the technical direction for ML training, serving, and observability.
β’ Act as the final escalation point for complex infrastructure issues related to GPU capacity, Spark tuning, and production incidents.
β’ Take ownership of and enhance the paved-road framework, which includes shared CI/CD, model-workflow scaffolding, and Databricks Asset Bundles.
β’ Lead the architecture for LLM endpoint serving across Databricks, AWS, Snowflake, and self-hosted deployments.
β’ Tackle challenges related to latency, cost, caching, evaluation, and PHI-safe routing for LLM serving.
β’ Manage standards and tools for MLflow, model registry, training image supply chain, and observability.
β’ Collaborate with Data & ML Platform, Data Science, App Dev, and Operations teams.
β’ Provide technical insights for vendor and platform selection processes.
β’ Mentor senior engineers and offer guidance to platform users.
β’ Write high-impact code and engage in hands-on Infrastructure-as-Code development.
β’ Over 10 years of software engineering experience, including a minimum of 3 years in designing, evolving, and operating enterprise-scale ML platforms in production.
β’ Strong technical judgment in ambiguous situations and a proven record of establishing standards and influencing colleagues.
β’ Production experience with Databricks and/or Amazon SageMaker.
β’ Familiarity with MLflow or a similar tracking and registry system.
β’ Experience with at least one core ML framework, such as PyTorch or TensorFlow.
β’ Proficiency in Java or a JVM equivalent, along with Python.
β’ Extensive experience with Apache Spark for large-scale data and distributed computing.
β’ In-depth knowledge of AWS, including networking, IAM, GPU compute, storage, and messaging services.
β’ Proficiency in Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads.
β’ Direct experience in serving LLMs in production, including cost management, evaluation harnesses, and secure handling of sensitive prompts and outputs.
β’ Regular use of Claude Code, Cursor, Copilot, or similar AI coding tools.
β’ Excellent written and verbal communication skills, particularly in asynchronous remote environments.
β’ This position is not eligible for employment sponsorship.
β’ Preferred: experience in technical leadership for healthcare or regulated-industry ML platforms; knowledge of Databricks Asset Bundles, Unity Catalog, Iceberg, Delta; specialized inference pipelines; GPU capacity planning; Kafka or Kinesis; clinical or safety-sensitive AI evaluation and red teaming; open-source ML infrastructure or production ML publications.
β’ Comprehensive total rewards compensation strategy.
β’ Post-offer health screenings and vaccinations as mandated by clients.
β’ Reasonable accommodations provided for individuals with physical and mental disabilities.
β’ Equal employment opportunity protections.
OPENDataJobs
Presidio
EasyLlama - HR & Compliance Training For Modern Teams
Coinbase
Get handpicked remote jobs straight to your inbox weekly.