
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United Kingdom.
β’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Spearhead architecture proposals, as well as high-level and low-level design plans, conduct technical reviews, and uphold the architecture baseline.
β’ Manage and lead a team of approximately twelve distributed engineers specializing in backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and provide input on growth and performance.
β’ Design, construct, and maintain a managed Slurm service for research users.
β’ Oversee Slurm components including controllers, accounting, partitions, login nodes, node onboarding, driver/CUDA standards, health detection, draining, autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal, including GPU and Network Operators, VM-based GPU isolation, upgrades, backup/recovery, and node replacement.
β’ Define the managed inference architecture encompassing multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Set up metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response initiatives, conduct post-incident reviews, and design a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and negotiate capacity.
β’ Complete the platform team and set a technical benchmark for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that build and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Practical experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and performing upgrades while jobs are on the system.
β’ Ideally, experience managing an HPC or GPU training cluster for a research community.
β’ Experience operating NVIDIA GPU fleets on bare metal, including managing the driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance testing.
β’ Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing issues related to multi-node NCCL performance.
β’ Strong knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
β’ Production experience with Kubernetes, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Experience with HPC storage solutions and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Familiarity with monitoring tools like Prometheus, Grafana, Loki, or their equivalents, as well as SLOs, incident response, and post-incident review processes.
β’ Proficiency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and influence architectural decisions.
β’ Experience launching a platform with real users, such as a multi-tenant IaaS/PaaS or research computing service, including aspects like resource isolation, quotas, usage metering, and user-facing API/CLI interfaces.
β’ Demonstrated ability in people management across different time zones, conducting cross-track reviews, drafting written architecture decisions, and effectively challenging partners or executives with well-reasoned arguments.
β’ Excellent command of written and spoken English.
β’ Must reside in a location between UTC and UTC+5:30.
β’ Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable experience with vLLM, SGLang, TensorRT-LLM, GPU isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and managing hardware-provider contracts with acceptance tests.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.