
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United Kingdom.
β’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Spearhead architecture proposals, oversee both high-level and low-level designs, perform technical reviews, and maintain the architectural baseline.
β’ Manage and lead a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Set engineering standards, conduct code and design reviews, manage release gates, hold one-on-one meetings, and provide feedback on growth and performance.
β’ Design, construct, and manage a Slurm service dedicated to research users.
β’ Oversee Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, as well as drainage and autohealing, storage visibility, identity, and isolation.
β’ Manage the Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement.
β’ Design the architecture for managed inference serving, focusing on multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, and SLOs throughout the control plane, GPU fleet, and application tiers.
β’ Lead incident response initiatives, conduct post-incident reviews, and design a sustainable on-call model.
β’ Act as the main technical liaison to infrastructure partners and vendors.
β’ Translate requirements into detailed specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.
β’ Complete the platform team and establish a high technical standard for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by others.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Proven experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and performing upgrades with jobs running on the system.
β’ Experience in operating an HPC or GPU training cluster for research users is highly preferred.
β’ Proficiency in GPU fleet operation on bare metal, including managing the NVIDIA driver and CUDA lifecycle, understanding Fabric Manager and NVSwitch behavior on SXM systems, DCGM-based health and utilization, MIG, and node burn-in and acceptance.
β’ Expertise in high-performance interconnects, particularly with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Strong knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-intensive workloads.
β’ Experience in production Kubernetes operation, encompassing control plane management, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design.
β’ Familiarity with HPC storage and data movement technologies such as VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across multiple nodes.
β’ Experience with observability and operations tools like Prometheus, Grafana and Loki or similar, SLOs, incident response, and post-incident reviews.
β’ Working proficiency in JavaScript and Node.js to effectively review control plane, CLI, and worker services and make informed architectural decisions.
β’ Experience deploying a platform with active users, such as a multi-tenant IaaS or PaaS or a research computing service.
β’ Proven ability to manage people across different time zones, conduct cross-track reviews, document architectural decisions, and effectively communicate those decisions to partners or executives.
β’ Excellent written and spoken English skills.
β’ Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers is desirable.
β’ Familiarity with modern serving stacks like vLLM, SGLang, or TensorRT-LLM is a plus.
β’ Desirable experience in VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and partnerships with hardware providers.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.