
Technical Lead – GPU Infrastructure
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants anywhere in the world.
• Take ownership of the complete platform architecture by creating architecture proposals, producing high-level and low-level designs, conducting reviews, and maintaining the baseline.
• Lead and manage a distributed team of approximately twelve engineers specializing in backend, frontend, DevOps, QA, and documentation.
• Define engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and provide input on growth and performance.
• Design, develop, and operate a managed Slurm service tailored for research users.
• Manage controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation.
• Oversee the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal.
• Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, along with upgrades, backup and recovery, and node replacement.
• Manage the architecture for inference, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
• Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application layers.
• Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.
• Act as the primary technical liaison to infrastructure partners and vendors.
• Convert requirements into documented specifications and acceptance tests, manage escalations, and assist with capacity planning and hardware procurement.
• Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.
• Complete the platform team and set a high technical standard for new engineers.
• Eight or more years of hands-on engineering experience.
• A minimum of three years in leadership roles overseeing teams that develop and manage infrastructure platforms relied upon by other teams.
• Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
• Practical experience running Slurm at scale, which includes slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are present on the system.
• Experience in managing an HPC or GPU training cluster for a research community is highly preferred.
• Proficient in operating NVIDIA GPU fleets on bare metal, covering NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.
• Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
• Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
• Experience with production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
• Knowledge of HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.
• Proficiency with Prometheus, Grafana, Loki or similar tools, including SLOs, incident response, and post-incident review processes.
• Working fluency in JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architectural decisions.
• Experience delivering a platform with real user engagement, such as a multi-tenant IaaS/PaaS or research computing service.
• People management experience across different time zones, cross-track review processes, written architectural decision-making, and communication with partners/executives.
• Excellent proficiency in written and spoken English.
• Must be located between UTC and UTC+5:30.
• Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
• Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
• Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.
• Fully remote work opportunity.
• Occasional travel to partner locations and team events.
Prothera
The Home Depot
Truelogic Software
Get handpicked remote jobs straight to your inbox weekly.