
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
β’ Take responsibility for the complete platform architecture by creating architecture proposals, as well as high-level and low-level designs, reviews, and maintaining the baseline.
β’ Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, manage release gates, hold one-on-ones, and provide insights for growth and performance.
β’ Design, implement, and operate a managed Slurm service tailored for research users.
β’ Manage Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation.
β’ Oversee the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal.
β’ Implement NVIDIA GPU Operator, Network Operator, VM-based GPU isolation using KubeVirt and VFIO, along with upgrades, backup and recovery, and node replacement.
β’ Create managed inference serving architecture, ensuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Set up metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and manage capacity.
β’ Complete the platform team and set the technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that build and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Extensive hands-on experience with Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.
β’ Preferred experience in operating HPC or GPU training clusters for research users.
β’ Familiarity with the lifecycle of NVIDIA drivers and CUDA.
β’ Experience with Fabric Manager and NVSwitch on SXM systems.
β’ Knowledge of DCGM-based health and utilization, MIG, node burn-in, and acceptance testing.
β’ Proficiency in InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.
β’ Experience with Linux kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
β’ Proven experience in production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Experience with HPC storage and data movement utilizing VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Familiarity with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and post-incident review processes.
β’ Working proficiency in JavaScript and Node.js adequate for reviewing and making architectural decisions.
β’ Experience with a deployed platform utilized by real users, such as multi-tenant IaaS, PaaS, or research computing services.
β’ Proven people management skills across time zones, cross-track reviews, written architectural decisions, and effective communication with partners and executives.
β’ Excellent proficiency in written and spoken English.
β’ Desirable experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU, KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.