
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Pakistan.
β’ Take ownership of the comprehensive architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Develop and sustain architecture proposals, high-level designs, and low-level designs through thorough review processes.
β’ Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Set engineering standards, conduct code and design reviews, manage release gates, hold one-on-one meetings, and provide input on growth and performance.
β’ Design, build, and manage a Slurm service tailored for research users.
β’ Oversee the Slurm controller and accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation.
β’ Manage the Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal.
β’ Handle NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, upgrades, backup and recovery, and node replacement.
β’ Define the architecture for managed inference, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, post-incident reviews, and maintain a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Transform requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.
β’ Complete the platform team and establish the technical standards for new engineers.
β’ A minimum of eight years of practical engineering experience.
β’ At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Hands-on experience managing Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and performing upgrades with jobs active on the system.
β’ Ideally, experience in operating an HPC or GPU training cluster for a research community.
β’ Experience managing NVIDIA GPU fleets on bare metal, encompassing NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
β’ Familiarity with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing performance issues with multi-node NCCL.
β’ Extensive knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
β’ Experience with production Kubernetes operation, covering control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Knowledge of HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Familiarity with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and post-incident reviews.
β’ Proficiency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architectural decisions.
β’ Experience deploying a platform with active users, such as a multi-tenant IaaS/PaaS or research computing service.
β’ Proven ability to manage people across various time zones, conduct cross-track reviews, document architecture decisions, and challenge partners or executives with rationale.
β’ Excellent command of written and spoken English.
β’ Fully remote role available for candidates located between UTC and UTC+5:30.
β’ Preferred experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Preferred experience with contemporary serving stacks like vLLM, SGLang, and TensorRT-LLM.
β’ Preferred experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, and GitOps.
β’ Preferred experience in GPU cloud, HPC center, AI lab platforms, peer-to-peer systems, distributed systems, or hardware-provider operations.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.