ML Cloud Infrastructure Engineer

atHavocAIRemoteUS flagUnited StatesFull-timeInfrastructure EngineerMid-levelSenior$150k – $175k/year

Posted 18 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Construct pipelines that transform raw multi-modal data into curated and versioned training datasets.

• Create reproducible training and evaluation workflows utilizing cloud computing and GPU resources.

• Develop and sustain model deployment infrastructure for packaging, serving, inference, versioning, and rollback.

• Implement capabilities for experiment tracking, dataset lineage, model versioning, and reproducible ML development.

• Manage data schema versioning and migration across pipelines, data lakes, and services.

• Design, build, and maintain scalable AWS infrastructure using Infrastructure as Code methodologies.

• Construct and manage Kubernetes/EKS workloads and containerized environments.

• Create self-service tools and paved paths for compute scheduling, storage, data access, training, and deployment.

• Enhance utilization, scalability, and cost efficiency across cloud and accelerator infrastructures.

• Develop evaluation frameworks and regression testing to ensure model quality, dataset integrity, and pipeline correctness.

• Create monitoring, logging, tracing, and observability systems for training jobs, data pipelines, and deployed models.

• Diagnose and address performance, scaling, reliability, and infrastructure bottlenecks.

• Maintain automation processes, testing, documentation, and operational readiness.

• Collaborate with Autonomy, Software, Data, Simulation, and Security teams on ML infrastructure from edge data capture to cloud training and model deployment.

• Contribute to CI/CD and release processes for models, datasets, and ML pipelines.

• Translate engineering requirements into scalable platform capabilities.

• Implement secure infrastructure practices including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data.

• Collaborate with security and infrastructure teams to meet operational and compliance requirements.

• Establish reliable and reproducible pipelines, scalable training and evaluation, self-service ML infrastructure, and improved observability, reliability, security, and cost efficiency within the first 12 months.


⛳️ Requirements

• Minimum of 3 years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related discipline.

• Proficient in programming with Python; experience in Go, C++, or other systems-oriented languages is preferred.

• Demonstrated experience in building and operating production services, APIs, data pipelines, developer platforms, or infrastructure.

• Practical experience with ML workflows, including dataset preparation, model training, evaluation, or deployment.

• Familiarity with cloud infrastructure, preferably AWS, and Infrastructure as Code practices.

• Hands-on experience with Kubernetes and containerized environments.

• Strong comprehension of production engineering fundamentals, such as reliability, observability, testing, automation, and maintainability.

• Capability to work efficiently across engineering disciplines and tackle ambiguous technical challenges with a strong sense of ownership.

• U.S. Citizenship is required.

• Ability to obtain and maintain a U.S. Government security clearance.

• Experience with MLOps and workflow platforms is preferred, including MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster.

• Preferred experience in GPU/accelerator scheduling, distributed training, or large-scale ML workloads.

• Familiarity with multi-modal datasets, including imagery, video, telemetry, sensor, or simulation data is preferred.

• Experience supporting autonomy, robotics, simulation, or real-time systems is preferred.

• Experience deploying ML models to edge or embedded environments is preferred.

• Familiarity with AWS GovCloud, GCP Assured Workloads, FedRAMP, or IL4/IL5 is preferred.


🏝️ Benefits

• 100% Employer paid Health, Dental, and Vision Insurance for you and your family.

• Life Insurance (Employer Paid).

• Opportunity to participate in the company's 401k program (with matching).

• Unlimited PTO policy with a mandated minimum of 2 weeks.

• Equity Package.

• Work / Home Office Stipend.

• Global Entry.

• 16 Weeks of Paid Parental Leave.

• Monthly Health and Wellness Stipend.

• Bonus opportunities available.

People also viewed

Sezzle17 hours ago

Principal Infrastructure Engineer

FR flagFrance OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
QuickNode ⚡22 hours ago

Senior Infrastructure Engineer, Core Systems

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

DE flagGermany OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

NL flagNetherlands OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

PL flagPoland OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
VALR3 days ago

Senior Infrastructure Engineer

ZA flagSouth Africa OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers