
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Pakistan.
β’ Take ownership of Cosmic AC's comprehensive platform architecture, including crafting architecture proposals, high-level and low-level designs, conducting reviews, and maintaining baselines.
β’ Supervise and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-one meetings, and provide feedback on growth and performance.
β’ Design, develop, and manage a Slurm service tailored for research users.
β’ Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation.
β’ Manage the Kubernetes cluster bootstrap and lifecycle on bare metal provided by partners.
β’ Operate the NVIDIA GPU Operator and Network Operator, implement VM-based GPU isolation using KubeVirt and VFIO, and handle upgrades, backup and recovery, and node replacements.
β’ Lead the architecture for managed inference, encompassing serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, and SLOs for the control plane, GPU fleet, and application layers.
β’ Direct incident response efforts, post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Transform requirements into detailed written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.
β’ Complete the platform team and set the technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that develop and maintain infrastructure platforms relied upon by other teams.
β’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Proven hands-on experience with Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and executing upgrades with jobs running.
β’ Ideally, experience operating an HPC or GPU training cluster for a research population.
β’ Familiarity with managing NVIDIA GPU fleets on bare metal, encompassing driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.
β’ Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
β’ Experience with production Kubernetes operations, covering control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.
β’ Background in HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Familiarity with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review.
β’ Proficient working knowledge of JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architectural decisions.
β’ Experience in delivering a multi-tenant IaaS, PaaS, or research computing service with isolation, quotas, usage metering, and user-facing API and CLI surfaces.
β’ Proven people management skills across time zones, cross-track review, written architectural decisions, and the ability to challenge partners or executives with sound reasoning.
β’ Excellent written and spoken English skills.
β’ Availability to work from a location between UTC and UTC+5:30.
β’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable: Knowledge of modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable: Familiarity with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platform experience, distributed systems, and hardware-provider partnership experience.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.