
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in India.
β’ Take charge of the Cosmic AC platform architecture from start to finish, including creating architecture proposals, high-level and low-level designs, conducting reviews, and maintaining the baseline.
β’ Lead and manage a distributed team of approximately twelve engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and provide insights for growth and performance.
β’ Design, develop, and manage a managed Slurm service tailored for research users.
β’ Oversee controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation.
β’ Manage the Kubernetes cluster bootstrap and lifecycle on bare metal provided by partners.
β’ Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, along with upgrades, backup and recovery, and node replacement.
β’ Oversee managed inference architecture, encompassing serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform requirements and negotiate limited capacity.
β’ Complete the platform team and set the technical standards for new hires.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that build and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Practical experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs in progress.
β’ Ideally, experience in operating an HPC or GPU training cluster for a research population is preferred.
β’ Experience with bare-metal NVIDIA GPU fleet operations, including the NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance testing.
β’ Knowledge of InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Expertise in Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience in production Kubernetes operations, covering control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Familiarity with HPC storage and data movement utilizing VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Experience with observability and operations tools such as Prometheus, Grafana, Loki or their equivalents, along with SLOs, incident response, and post-incident review processes.
β’ Working proficiency in JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architectural decisions.
β’ Proven experience in delivering a platform for real users, such as multi-tenant IaaS/PaaS or research computing services, including aspects of isolation, quotas, usage metering, APIs, and CLIs.
β’ Skills in people management across time zones, conducting cross-track reviews, documenting architectural decisions in writing, and the ability to decline partner or executive requests with clear reasoning.
β’ Excellent written and verbal communication skills in English.
β’ Ability to work remotely from a location within the UTC to UTC+5:30 time zones.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.