
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in India.
β’ Take ownership of Cosmic AC's comprehensive platform architecture by creating architecture proposals, developing high-level and low-level designs, conducting reviews, and maintaining the baseline.
β’ Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-ones, and provide insights on growth and performance.
β’ Design, construct, and manage a Slurm service tailored for research users.
β’ Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, as well as stalled-job and node-health detection, draining, auto-healing, storage visibility, identity, and isolation.
β’ Take charge of the Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal.
β’ Implement NVIDIA GPU Operator and Network Operator, along with VM-based GPU isolation using KubeVirt and VFIO, upgrades, backup and recovery, and node replacement.
β’ Design a managed inference architecture featuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Manage metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, conduct post-incident reviews, and establish a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Translate requirements into detailed written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and negotiate capacity.
β’ Complete the platform team and set high technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams responsible for building and operating infrastructure platforms relied upon by other teams.
β’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs in the system.
β’ Experience operating HPC or GPU training clusters for research users is highly desirable.
β’ Experience managing NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.
β’ Proficiency with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Strong Linux systems expertise, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.
β’ Experience in production Kubernetes operations, focusing on control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.
β’ Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Experience with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and conducting post-incident reviews.
β’ Working proficiency in JavaScript and Node.js sufficient to evaluate and make architectural decisions regarding control plane, CLI, and worker services.
β’ Experience delivering a deployed platform with actual users, such as multi-tenant IaaS/PaaS or research computing services.
β’ Capable of managing people across different time zones, conducting cross-track reviews, documenting architectural decisions, and effectively challenging partners or executives with rationale.
β’ Excellent written and spoken English skills.
β’ Must reside between UTC and UTC+5:30.
β’ Willingness to travel occasionally to partner sites and team events.
β’ Preferred experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class auto-healing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, peer-to-peer or distributed systems, and hardware-provider partnerships.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
β’ Global, distributed team environment.
β’ Opportunities for line-management and growth/performance development.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.