
Tech Lead Manager, GPU Cluster Infrastructure
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in California.
β’ Establish the technical vision for the platform and take responsibility for its roadmap.
β’ Determine which systems to operate, manage scheduling and storage across providers, and identify key metrics for measurement.
β’ Lead the design of architecture, scheduling, and storage systems.
β’ Troubleshoot failures across various layers, including node health, GPU and fabric issues, and multi-node job stalls.
β’ Recruit and develop a small team of senior engineers.
β’ Prioritize tasks, define ownership, scope projects, and offer regular feedback and coaching.
β’ Set the security framework for a shared cluster where AI agents conduct experiments, focusing on identity and access, workload isolation, and sandboxing.
β’ Outline operational procedures, including on-call rotations, incident response, postmortems, fault tolerance, and observability.
β’ Engage in operational practices alongside the team.
β’ Act as the escalation point for research teams and infrastructure providers.
β’ Transform recurring issues into effective platform solutions.
β’ Collaborate directly with researchers and engineers to ensure large-scale experiments are efficient and resilient.
β’ Proven experience in leading engineers as a manager, tech lead, or project lead, encompassing the setting of technical direction, project scoping, and feedback provision.
β’ A strong desire to manage people directly; previous experience with direct reports is not a prerequisite.
β’ Over 5 years of experience in systems or infrastructure engineering within a production Linux environment.
β’ Experience in managing GPU, HPC, or large-scale batch processing platforms.
β’ Proven track record of owning systems from the design phase through to operational execution.
β’ In-depth knowledge in at least one area such as scheduling, storage, networking, security, or GPU systems, along with sufficient breadth to evaluate designs in other areas.
β’ Experience managing production Kubernetes for GPU workloads with a batch processing layer like Slurm, Kueue, or Volcano, including management of quotas, priority and preemption, and node health.
β’ Familiarity with infrastructure as code and observability practices for a production fleet.
β’ Proficiency with tools such as Terraform or Ansible, Helm, ArgoCD, and Prometheus, or their equivalents.
β’ Strong programming skills in languages like Python, Go, Rust, C++, or another commonly utilized infrastructure language.
β’ Ability to produce clear technical documentation for engineers, researchers, and service providers.
β’ Additional desirable skills may include expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, multi-provider platforms, and building teams from the ground up.
β’ Health insurance with 94% of the premium covered by the organization, effective within one month of the start date.
β’ 401(k) plan featuring up to a 2% match.
β’ Accrual of 25 days of paid time off per year, calculated weekly.
β’ Up to 10 days of paid sick leave annually.
β’ Paid leave for bereavement, family, medical, and pregnancy disability reasons.
β’ Provision of a work computer and WFH stipend for eligible employees.
β’ Catered lunches and dinners provided on workdays at the Berkeley office.
β’ Visa sponsorship available for employees working in-person.
β’ Paid work trial lasting up to one week.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.