
Technical Lead – GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the complete platform architecture by overseeing proposals, high-level and low-level designs, reviews, and baseline maintenance.
• Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation domains.
• Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide insights on growth and performance.
• Design, build, and operate a managed Slurm service for research users.
• Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health detection, autohealing, storage visibility, identity, and isolation.
• Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal, including NVIDIA GPU and Network Operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement.
• Design managed inference architecture that includes multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
• Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and a sustainable on-call model.
• Act as the primary technical liaison to infrastructure partners and vendors.
• Convert requirements into written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
• Collaborate with research, model-training, and product teams to translate workloads into platform requirements and allocate limited capacity.
• Complete the platform team and set high technical standards for new engineers.
• A minimum of eight years of hands-on engineering experience.
• At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by other teams.
• Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
• Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.
• Preferred experience in operating an HPC or GPU training cluster for a research audience.
• Experience in operating NVIDIA GPU fleets on bare metal, covering driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.
• Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.
• Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
• Experience with production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
• Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
• Experience with Prometheus, Grafana, Loki or similar tools, along with SLOs, incident response, and post-incident reviews.
• Proficiency in JavaScript and Node.js sufficient to review and make architectural decisions regarding control plane, CLI, and worker services.
• Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.
• People management experience across time zones, with skills in cross-track review, written architecture decisions, and technical judgment in collaboration with partners and executives.
• Exceptional written and verbal communication skills in English.
• Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
• Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
• Desirable experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships.
• Fully remote work.
• Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.