Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in United Kingdom.

πŸ“‹ Description

β€’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.

β€’ Spearhead architecture proposals, oversee both high-level and low-level designs, perform technical reviews, and maintain the architectural baseline.

β€’ Manage and lead a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.

β€’ Set engineering standards, conduct code and design reviews, manage release gates, hold one-on-one meetings, and provide feedback on growth and performance.

β€’ Design, construct, and manage a Slurm service dedicated to research users.

β€’ Oversee Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, as well as drainage and autohealing, storage visibility, identity, and isolation.

β€’ Manage the Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement.

β€’ Design the architecture for managed inference serving, focusing on multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

β€’ Establish metrics, logging, alerting, and SLOs throughout the control plane, GPU fleet, and application tiers.

β€’ Lead incident response initiatives, conduct post-incident reviews, and design a sustainable on-call model.

β€’ Act as the main technical liaison to infrastructure partners and vendors.

β€’ Translate requirements into detailed specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing.

β€’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.

β€’ Complete the platform team and establish a high technical standard for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience.

β€’ At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by others.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Proven experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and performing upgrades with jobs running on the system.

β€’ Experience in operating an HPC or GPU training cluster for research users is highly preferred.

β€’ Proficiency in GPU fleet operation on bare metal, including managing the NVIDIA driver and CUDA lifecycle, understanding Fabric Manager and NVSwitch behavior on SXM systems, DCGM-based health and utilization, MIG, and node burn-in and acceptance.

β€’ Expertise in high-performance interconnects, particularly with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

β€’ Strong knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-intensive workloads.

β€’ Experience in production Kubernetes operation, encompassing control plane management, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design.

β€’ Familiarity with HPC storage and data movement technologies such as VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across multiple nodes.

β€’ Experience with observability and operations tools like Prometheus, Grafana and Loki or similar, SLOs, incident response, and post-incident reviews.

β€’ Working proficiency in JavaScript and Node.js to effectively review control plane, CLI, and worker services and make informed architectural decisions.

β€’ Experience deploying a platform with active users, such as a multi-tenant IaaS or PaaS or a research computing service.

β€’ Proven ability to manage people across different time zones, conduct cross-track reviews, document architectural decisions, and effectively communicate those decisions to partners or executives.

β€’ Excellent written and spoken English skills.

β€’ Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers is desirable.

β€’ Familiarity with modern serving stacks like vLLM, SGLang, or TensorRT-LLM is a plus.

β€’ Desirable experience in VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

β€’ Fully remote work arrangement.

β€’ Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers