
ML Cloud Infrastructure Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in United States.
• Construct pipelines that transform raw multi-modal data into curated and versioned training datasets.
• Create reproducible training and evaluation workflows utilizing cloud computing and GPU resources.
• Develop and sustain model deployment infrastructure for packaging, serving, inference, versioning, and rollback.
• Implement capabilities for experiment tracking, dataset lineage, model versioning, and reproducible ML development.
• Manage data schema versioning and migration across pipelines, data lakes, and services.
• Design, build, and maintain scalable AWS infrastructure using Infrastructure as Code methodologies.
• Construct and manage Kubernetes/EKS workloads and containerized environments.
• Create self-service tools and paved paths for compute scheduling, storage, data access, training, and deployment.
• Enhance utilization, scalability, and cost efficiency across cloud and accelerator infrastructures.
• Develop evaluation frameworks and regression testing to ensure model quality, dataset integrity, and pipeline correctness.
• Create monitoring, logging, tracing, and observability systems for training jobs, data pipelines, and deployed models.
• Diagnose and address performance, scaling, reliability, and infrastructure bottlenecks.
• Maintain automation processes, testing, documentation, and operational readiness.
• Collaborate with Autonomy, Software, Data, Simulation, and Security teams on ML infrastructure from edge data capture to cloud training and model deployment.
• Contribute to CI/CD and release processes for models, datasets, and ML pipelines.
• Translate engineering requirements into scalable platform capabilities.
• Implement secure infrastructure practices including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data.
• Collaborate with security and infrastructure teams to meet operational and compliance requirements.
• Establish reliable and reproducible pipelines, scalable training and evaluation, self-service ML infrastructure, and improved observability, reliability, security, and cost efficiency within the first 12 months.
• Minimum of 3 years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related discipline.
• Proficient in programming with Python; experience in Go, C++, or other systems-oriented languages is preferred.
• Demonstrated experience in building and operating production services, APIs, data pipelines, developer platforms, or infrastructure.
• Practical experience with ML workflows, including dataset preparation, model training, evaluation, or deployment.
• Familiarity with cloud infrastructure, preferably AWS, and Infrastructure as Code practices.
• Hands-on experience with Kubernetes and containerized environments.
• Strong comprehension of production engineering fundamentals, such as reliability, observability, testing, automation, and maintainability.
• Capability to work efficiently across engineering disciplines and tackle ambiguous technical challenges with a strong sense of ownership.
• U.S. Citizenship is required.
• Ability to obtain and maintain a U.S. Government security clearance.
• Experience with MLOps and workflow platforms is preferred, including MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster.
• Preferred experience in GPU/accelerator scheduling, distributed training, or large-scale ML workloads.
• Familiarity with multi-modal datasets, including imagery, video, telemetry, sensor, or simulation data is preferred.
• Experience supporting autonomy, robotics, simulation, or real-time systems is preferred.
• Experience deploying ML models to edge or embedded environments is preferred.
• Familiarity with AWS GovCloud, GCP Assured Workloads, FedRAMP, or IL4/IL5 is preferred.
• 100% Employer paid Health, Dental, and Vision Insurance for you and your family.
• Life Insurance (Employer Paid).
• Opportunity to participate in the company's 401k program (with matching).
• Unlimited PTO policy with a mandated minimum of 2 weeks.
• Equity Package.
• Work / Home Office Stipend.
• Global Entry.
• 16 Weeks of Paid Parental Leave.
• Monthly Health and Wellness Stipend.
• Bonus opportunities available.
Sezzle
QuickNode ⚡
Sezzle
Sezzle
Get handpicked remote jobs straight to your inbox weekly.