Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the complete platform architecture by overseeing proposals, high-level and low-level designs, reviews, and baseline maintenance.

• Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation domains.

• Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide insights on growth and performance.

• Design, build, and operate a managed Slurm service for research users.

• Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health detection, autohealing, storage visibility, identity, and isolation.

• Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal, including NVIDIA GPU and Network Operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement.

• Design managed inference architecture that includes multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

• Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and a sustainable on-call model.

• Act as the primary technical liaison to infrastructure partners and vendors.

• Convert requirements into written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.

• Collaborate with research, model-training, and product teams to translate workloads into platform requirements and allocate limited capacity.

• Complete the platform team and set high technical standards for new engineers.


⛳️ Requirements

• A minimum of eight years of hands-on engineering experience.

• At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by other teams.

• Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

• Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.

• Preferred experience in operating an HPC or GPU training cluster for a research audience.

• Experience in operating NVIDIA GPU fleets on bare metal, covering driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.

• Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.

• Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.

• Experience with production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.

• Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.

• Experience with Prometheus, Grafana, Loki or similar tools, along with SLOs, incident response, and post-incident reviews.

• Proficiency in JavaScript and Node.js sufficient to review and make architectural decisions regarding control plane, CLI, and worker services.

• Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.

• People management experience across time zones, with skills in cross-track review, written architecture decisions, and technical judgment in collaboration with partners and executives.

• Exceptional written and verbal communication skills in English.

• Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

• Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.

• Desirable experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships.


🏝️ Benefits

• Fully remote work.

• Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers