
Senior Software Engineer, GPU Cluster Infrastructure
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in United States, +2 more locations.
β’ Manage the day-to-day operations of the Kubernetes GPU fleet, which includes overseeing node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.
β’ Take ownership of batch scheduling and multi-tenancy, including handling queues, quotas, priorities, preemption, gang scheduling, and ensuring fair sharing across research teams.
β’ Design and manage storage solutions under the fleet, incorporating high-performance shared filesystems, object storage tiers, quotas, and backup systems.
β’ Ensure fault tolerance in multi-node training runs by managing node health, automating draining, troubleshooting NCCL and fabric issues, monitoring stragglers and flaky GPUs, and implementing checkpoint and restart strategies.
β’ Strengthen the platform's security through identity and access management, network policies, secret handling, workload isolation, and sandboxing for AI agents.
β’ Facilitate the integration of new capacity by testing providers on fabric, NCCL, and storage throughput, ensuring compliance with SLAs, and incorporating new clusters using infrastructure as code.
β’ Collaborate closely with research teams to address infrastructure challenges and convert recurring issues into long-term platform solutions.
β’ Participate in the on-call rotation, maintain runbooks, and conduct postmortems.
β’ Work alongside researchers and engineers to ensure large-scale experiments are both efficient and resilient.
β’ A minimum of 3 years of experience in systems or infrastructure engineering focused on production Linux environments, particularly with GPU, HPC, or large-scale batch platforms.
β’ Proven experience in owning at least one system from its design phase through to operational deployment.
β’ Hands-on experience with production Kubernetes for GPU workloads, utilizing a batch layer such as Slurm, Kueue, Volcano, or a comparable system.
β’ Familiarity with managing quotas, priorities, preemption, and node health monitoring.
β’ Experience in infrastructure as code practices and observability metrics for a production fleet.
β’ Proficiency in Terraform or Ansible.
β’ Experience with deployment tools such as Helm and ArgoCD.
β’ Knowledge of monitoring solutions like Prometheus or equivalent technologies.
β’ Strong programming capabilities in at least one infrastructure language, such as Python, Go, Rust, or C++.
β’ Ability to communicate clearly with engineers, researchers, and service providers.
β’ Additional relevant expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, or multi-provider platforms is a plus.
β’ Health Insurance - 94% of the insurance premium covered by the organization, effective within one month of your start date.
β’ Retirement - 401(k) plan with a matching contribution of up to 2%.
β’ 25 days of Paid Time Off per year, accrued on a weekly basis.
β’ Up to 10 days of paid sick leave annually.
β’ Paid leave for bereavement, family, medical, and pregnancy disability purposes.
β’ Work-from-home stipend and a work computer provided for eligible employees.
β’ Catered lunches and dinners on workdays at the Berkeley office.
β’ Reimbursement for work-related travel and equipment expenses.
β’ Visa sponsorship available for in-person employees.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.