Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in India.

πŸ“‹ Description

β€’ Take ownership of Cosmic AC's comprehensive platform architecture by creating architecture proposals, developing high-level and low-level designs, conducting reviews, and maintaining the baseline.

β€’ Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.

β€’ Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-ones, and provide insights on growth and performance.

β€’ Design, construct, and manage a Slurm service tailored for research users.

β€’ Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, as well as stalled-job and node-health detection, draining, auto-healing, storage visibility, identity, and isolation.

β€’ Take charge of the Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal.

β€’ Implement NVIDIA GPU Operator and Network Operator, along with VM-based GPU isolation using KubeVirt and VFIO, upgrades, backup and recovery, and node replacement.

β€’ Design a managed inference architecture featuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

β€’ Manage metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.

β€’ Lead incident response efforts, conduct post-incident reviews, and establish a sustainable on-call model.

β€’ Act as the primary technical liaison to infrastructure partners and vendors.

β€’ Translate requirements into detailed written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing.

β€’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and negotiate capacity.

β€’ Complete the platform team and set high technical standards for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience.

β€’ At least three years of experience leading teams responsible for building and operating infrastructure platforms relied upon by other teams.

β€’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs in the system.

β€’ Experience operating HPC or GPU training clusters for research users is highly desirable.

β€’ Experience managing NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.

β€’ Proficiency with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

β€’ Strong Linux systems expertise, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning.

β€’ Experience in production Kubernetes operations, focusing on control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.

β€’ Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.

β€’ Experience with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and conducting post-incident reviews.

β€’ Working proficiency in JavaScript and Node.js sufficient to evaluate and make architectural decisions regarding control plane, CLI, and worker services.

β€’ Experience delivering a deployed platform with actual users, such as multi-tenant IaaS/PaaS or research computing services.

β€’ Capable of managing people across different time zones, conducting cross-track reviews, documenting architectural decisions, and effectively challenging partners or executives with rationale.

β€’ Excellent written and spoken English skills.

β€’ Must reside between UTC and UTC+5:30.

β€’ Willingness to travel occasionally to partner sites and team events.

β€’ Preferred experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class auto-healing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, peer-to-peer or distributed systems, and hardware-provider partnerships.


🏝️ Benefits

β€’ Fully remote work arrangement.

β€’ Occasional travel to partner sites and team events.

β€’ Global, distributed team environment.

β€’ Opportunities for line-management and growth/performance development.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers