Tech Lead Manager, GPU Cluster Infrastructure

Posted Sep 12

This is a fully remote position, open to applicants in California.

πŸ“‹ Description

β€’ Establish the technical vision for the platform and take responsibility for its roadmap.

β€’ Determine which systems to operate, manage scheduling and storage across providers, and identify key metrics for measurement.

β€’ Lead the design of architecture, scheduling, and storage systems.

β€’ Troubleshoot failures across various layers, including node health, GPU and fabric issues, and multi-node job stalls.

β€’ Recruit and develop a small team of senior engineers.

β€’ Prioritize tasks, define ownership, scope projects, and offer regular feedback and coaching.

β€’ Set the security framework for a shared cluster where AI agents conduct experiments, focusing on identity and access, workload isolation, and sandboxing.

β€’ Outline operational procedures, including on-call rotations, incident response, postmortems, fault tolerance, and observability.

β€’ Engage in operational practices alongside the team.

β€’ Act as the escalation point for research teams and infrastructure providers.

β€’ Transform recurring issues into effective platform solutions.

β€’ Collaborate directly with researchers and engineers to ensure large-scale experiments are efficient and resilient.


⛳️ Requirements

β€’ Proven experience in leading engineers as a manager, tech lead, or project lead, encompassing the setting of technical direction, project scoping, and feedback provision.

β€’ A strong desire to manage people directly; previous experience with direct reports is not a prerequisite.

β€’ Over 5 years of experience in systems or infrastructure engineering within a production Linux environment.

β€’ Experience in managing GPU, HPC, or large-scale batch processing platforms.

β€’ Proven track record of owning systems from the design phase through to operational execution.

β€’ In-depth knowledge in at least one area such as scheduling, storage, networking, security, or GPU systems, along with sufficient breadth to evaluate designs in other areas.

β€’ Experience managing production Kubernetes for GPU workloads with a batch processing layer like Slurm, Kueue, or Volcano, including management of quotas, priority and preemption, and node health.

β€’ Familiarity with infrastructure as code and observability practices for a production fleet.

β€’ Proficiency with tools such as Terraform or Ansible, Helm, ArgoCD, and Prometheus, or their equivalents.

β€’ Strong programming skills in languages like Python, Go, Rust, C++, or another commonly utilized infrastructure language.

β€’ Ability to produce clear technical documentation for engineers, researchers, and service providers.

β€’ Additional desirable skills may include expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, multi-provider platforms, and building teams from the ground up.


🏝️ Benefits

β€’ Health insurance with 94% of the premium covered by the organization, effective within one month of the start date.

β€’ 401(k) plan featuring up to a 2% match.

β€’ Accrual of 25 days of paid time off per year, calculated weekly.

β€’ Up to 10 days of paid sick leave annually.

β€’ Paid leave for bereavement, family, medical, and pregnancy disability reasons.

β€’ Provision of a work computer and WFH stipend for eligible employees.

β€’ Catered lunches and dinners provided on workdays at the Berkeley office.

β€’ Visa sponsorship available for employees working in-person.

β€’ Paid work trial lasting up to one week.

People also viewed

Verwaltungscloud.SH GmbH1 day ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS1 day ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers