
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Spain.
β’ Take ownership of the comprehensive architecture for Cosmic AC, which includes making architecture proposals, developing high-level and low-level designs, conducting reviews, and managing the baseline.
β’ Lead and manage a distributed team of around twelve engineers across various domains such as backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and offer insights on growth and performance.
β’ Design, implement, and manage a Slurm service for research users.
β’ Oversee Slurm controller and accounting, manage partitions, login nodes, node onboarding, driver and CUDA baselines, as well as stalled-job and node-health detection, draining and autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal.
β’ Implement NVIDIA GPU Operator and Network Operator, manage VM-based GPU isolation with KubeVirt and VFIO, handle upgrades, backup and recovery, and node replacement.
β’ Lead the architecture for managed inference, focusing on serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application layers.
β’ Direct incident response efforts, post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison with infrastructure partners and vendors.
β’ Convert requirements into detailed specifications and acceptance tests, manage escalations, and provide technical guidance for capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform needs and facilitate capacity brokering.
β’ Complete the platform team and elevate the technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience, with at least three years in leadership roles managing teams that develop and sustain critical infrastructure platforms.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Direct experience running Slurm at scale, including managing slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and system upgrades with active jobs.
β’ Experience in operating an HPC or GPU training cluster for a research community is highly preferred.
β’ Proven experience managing NVIDIA GPU fleets on bare metal, including lifecycle management of NVIDIA driver and CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
β’ Proficiency with InfiniBand, subnet configuration, RDMA, SR-IOV, and resolving multi-node NCCL performance issues.
β’ Extensive knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization techniques.
β’ Experience in production Kubernetes operations, including handling control plane upgrades, CNI, CSI, operators, custom controllers, and designing for multi-tenancy.
β’ Familiarity with HPC storage solutions and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Experience with monitoring tools like Prometheus, Grafana, Loki, or similar, including SLOs, incident response protocols, and post-incident review processes.
β’ Proficient in JavaScript and Node.js, with the ability to review a control plane, CLI, and worker services and make architectural decisions.
β’ Proven experience in delivering platforms with real users, such as multi-tenant IaaS/PaaS or research computing services, including aspects like resource isolation, quotas, usage monitoring, and user-facing API/CLI interfaces.
β’ Effective people management skills across time zones, including cross-track reviews, documented architectural decisions, and the ability to communicate effectively with partners or executives.
β’ Excellent proficiency in both written and spoken English.
β’ Available to work within UTC and UTC+5:30 time zones to ensure overlap with Europe and India.
β’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable: Knowledge of VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.