Senior Infrastructure Engineer – AI/ML Platform

atOpenTeamsRemoteUS flagUnited StatesFull-timeInfrastructure EngineerSenior$145k – $250k/year

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop and manage the Kubernetes platform that facilitates AI test and evaluation frameworks.

• Execute GPU scheduling, workload orchestration, resource management, and multi-tenant isolation for evaluation teams.

• Create infrastructure-as-code, GitOps workflows, and automated deployment pipelines to ensure reproducible platforms.

• Build reusable and modular infrastructure components for platforms that are independently owned and operated.

• Contribute to Nebari and other open-source initiatives related to Kubernetes, infrastructure, and MLOps.

• Take responsibility for platform reliability, capacity planning, upgrade strategies, failure-mode analysis, backup and recovery, and operational readiness.

• Design and implement observability, monitoring, logging, tracing, and alerting systems for large-scale AI/ML workloads.

• Create operational runbooks and documentation for deployment, operation, and troubleshooting tasks.

• Deploy, configure, and secure infrastructure within Government environments that are secure, restricted, or have limited connectivity.

• Support security authorization and compliance through infrastructure documentation, secure configurations, control evidence, and repeatable deployment processes.

• Integrate automated security tools for container scanning, static and dynamic analysis, artifact signing, and policy enforcement.

• Collaborate with Government stakeholders, security experts, software engineers, and ML engineers.

• Provide technical leadership, contribute to engineering standards, and mentor junior team members.

• Work effectively within a remote and distributed team using asynchronous communication.


⛳️ Requirements

• U.S. citizenship and the ability to obtain and maintain a Secret security clearance.

• Over 6 years of practical experience in infrastructure, platform, DevOps, or site reliability engineering supporting production systems.

• In-depth knowledge of scalability, reliability, observability, security, and automation practices.

• Practical experience with Kubernetes, including workload scheduling, resource management, and multi-tenant environments.

• Familiarity with automated security tools, including container scanning, SAST/DAST, artifact signing, and policy enforcement.

• Proficient in infrastructure-as-code tools such as Terraform, OpenTofu, Pulumi, or similar technologies.

• Experience with at least one major cloud platform—AWS, Azure, or Google Cloud—covering networking, security, storage, and compute services.

• Practical experience in implementing monitoring and observability using OpenTelemetry, Prometheus, Grafana, or equivalent technologies.

• Strong programming or automation abilities in Python, Go, or a comparable language.

• Familiarity with CI/CD practices, GitOps workflows, and infrastructure automation.

• Experience creating maintainable operational documentation, deployment procedures, and runbooks.

• Experience leading technical initiatives, establishing engineering practices, or mentoring fellow engineers.

• Ability to work independently while effectively collaborating in a remote, distributed team environment.

• Capability to provide and receive constructive technical feedback.


🏝️ Benefits

• Medical, Dental & Vision – 100% covered for employees, 75% for dependents.

• 401(k) Match – Up to 5% with full vesting after 2 years.

• Unlimited PTO – With a required minimum of 15 days off each year.

• Fully Remote Setup – Includes up to $3,000 equipment reimbursement.

• Continuous Education – Includes up to $500 reimbursement.

• Disability & Life Insurance – 100% employer-covered.

• HSA & FSA Options – With monthly HSA contributions from OpenTeams.

• Career framework that offers a pathway and recognition for increased impact.

• Opportunities for open-source contributions.

• Flexible asynchronous communication within a distributed team.

• Commitment to equal opportunity, diversity, equity, inclusion, and belonging.

People also viewed

Hitachi Solutions America1 day ago

Senior Azure Infrastructure Architect

DE flagGermany, +1 more countryFull-timeInfrastructure Engineer$140k – $180k/year
ApplyView job
Nava1 day ago

Senior Infrastructure Engineer, Azure

US flagAlabama, +29 more statesFull-timeInfrastructure Engineer$135.9k – $153k/year
ApplyView job
Kalepa1 day ago

Staff Platform Engineer – Infrastructure

US flagUnited States OnlyFull-timeInfrastructure Engineer$220k – $300k/year
ApplyView job
ElevenLabs1 day ago

HPC Infrastructure Engineer – GPU Clusters

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
Activision Blizzard1 day ago

Senior Azure Infrastructure Engineer

US flagCalifornia OnlyFull-timeInfrastructure Engineer$102.8k – $190.2k/year
ApplyView job
Activision1 day ago

Senior Azure Infrastructure Engineer

US flagCalifornia OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers