Senior Software Engineer, GPU Cluster Infrastructure

Posted Sep 12

This is a fully remote position, open to applicants in United States, +2 more locations.

πŸ“‹ Description

β€’ Manage the day-to-day operations of the Kubernetes GPU fleet, which includes overseeing node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.

β€’ Take ownership of batch scheduling and multi-tenancy, including handling queues, quotas, priorities, preemption, gang scheduling, and ensuring fair sharing across research teams.

β€’ Design and manage storage solutions under the fleet, incorporating high-performance shared filesystems, object storage tiers, quotas, and backup systems.

β€’ Ensure fault tolerance in multi-node training runs by managing node health, automating draining, troubleshooting NCCL and fabric issues, monitoring stragglers and flaky GPUs, and implementing checkpoint and restart strategies.

β€’ Strengthen the platform's security through identity and access management, network policies, secret handling, workload isolation, and sandboxing for AI agents.

β€’ Facilitate the integration of new capacity by testing providers on fabric, NCCL, and storage throughput, ensuring compliance with SLAs, and incorporating new clusters using infrastructure as code.

β€’ Collaborate closely with research teams to address infrastructure challenges and convert recurring issues into long-term platform solutions.

β€’ Participate in the on-call rotation, maintain runbooks, and conduct postmortems.

β€’ Work alongside researchers and engineers to ensure large-scale experiments are both efficient and resilient.


⛳️ Requirements

β€’ A minimum of 3 years of experience in systems or infrastructure engineering focused on production Linux environments, particularly with GPU, HPC, or large-scale batch platforms.

β€’ Proven experience in owning at least one system from its design phase through to operational deployment.

β€’ Hands-on experience with production Kubernetes for GPU workloads, utilizing a batch layer such as Slurm, Kueue, Volcano, or a comparable system.

β€’ Familiarity with managing quotas, priorities, preemption, and node health monitoring.

β€’ Experience in infrastructure as code practices and observability metrics for a production fleet.

β€’ Proficiency in Terraform or Ansible.

β€’ Experience with deployment tools such as Helm and ArgoCD.

β€’ Knowledge of monitoring solutions like Prometheus or equivalent technologies.

β€’ Strong programming capabilities in at least one infrastructure language, such as Python, Go, Rust, or C++.

β€’ Ability to communicate clearly with engineers, researchers, and service providers.

β€’ Additional relevant expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, or multi-provider platforms is a plus.


🏝️ Benefits

β€’ Health Insurance - 94% of the insurance premium covered by the organization, effective within one month of your start date.

β€’ Retirement - 401(k) plan with a matching contribution of up to 2%.

β€’ 25 days of Paid Time Off per year, accrued on a weekly basis.

β€’ Up to 10 days of paid sick leave annually.

β€’ Paid leave for bereavement, family, medical, and pregnancy disability purposes.

β€’ Work-from-home stipend and a work computer provided for eligible employees.

β€’ Catered lunches and dinners on workdays at the Berkeley office.

β€’ Reimbursement for work-related travel and equipment expenses.

β€’ Visa sponsorship available for in-person employees.

People also viewed

Verwaltungscloud.SH GmbH1 day ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS1 day ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers