Technical Lead – GPU Infrastructure

Posted 6 days ago

This is a fully remote position, open to applicants in Spain.

πŸ“‹ Description

β€’ Take ownership of the comprehensive architecture for Cosmic AC, which includes making architecture proposals, developing high-level and low-level designs, conducting reviews, and managing the baseline.

β€’ Lead and manage a distributed team of around twelve engineers across various domains such as backend, frontend, DevOps, QA, and documentation.

β€’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and offer insights on growth and performance.

β€’ Design, implement, and manage a Slurm service for research users.

β€’ Oversee Slurm controller and accounting, manage partitions, login nodes, node onboarding, driver and CUDA baselines, as well as stalled-job and node-health detection, draining and autohealing, storage visibility, identity, and isolation.

β€’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal.

β€’ Implement NVIDIA GPU Operator and Network Operator, manage VM-based GPU isolation with KubeVirt and VFIO, handle upgrades, backup and recovery, and node replacement.

β€’ Lead the architecture for managed inference, focusing on serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

β€’ Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application layers.

β€’ Direct incident response efforts, post-incident reviews, and develop a sustainable on-call model.

β€’ Act as the primary technical liaison with infrastructure partners and vendors.

β€’ Convert requirements into detailed specifications and acceptance tests, manage escalations, and provide technical guidance for capacity planning and hardware sourcing.

β€’ Collaborate with research, model-training, and product teams to translate workloads into platform needs and facilitate capacity brokering.

β€’ Complete the platform team and elevate the technical standards for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience, with at least three years in leadership roles managing teams that develop and sustain critical infrastructure platforms.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Direct experience running Slurm at scale, including managing slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and system upgrades with active jobs.

β€’ Experience in operating an HPC or GPU training cluster for a research community is highly preferred.

β€’ Proven experience managing NVIDIA GPU fleets on bare metal, including lifecycle management of NVIDIA driver and CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.

β€’ Proficiency with InfiniBand, subnet configuration, RDMA, SR-IOV, and resolving multi-node NCCL performance issues.

β€’ Extensive knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization techniques.

β€’ Experience in production Kubernetes operations, including handling control plane upgrades, CNI, CSI, operators, custom controllers, and designing for multi-tenancy.

β€’ Familiarity with HPC storage solutions and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.

β€’ Experience with monitoring tools like Prometheus, Grafana, Loki, or similar, including SLOs, incident response protocols, and post-incident review processes.

β€’ Proficient in JavaScript and Node.js, with the ability to review a control plane, CLI, and worker services and make architectural decisions.

β€’ Proven experience in delivering platforms with real users, such as multi-tenant IaaS/PaaS or research computing services, including aspects like resource isolation, quotas, usage monitoring, and user-facing API/CLI interfaces.

β€’ Effective people management skills across time zones, including cross-track reviews, documented architectural decisions, and the ability to communicate effectively with partners or executives.

β€’ Excellent proficiency in both written and spoken English.

β€’ Available to work within UTC and UTC+5:30 time zones to ensure overlap with Europe and India.

β€’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

β€’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.

β€’ Desirable: Knowledge of VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

β€’ Fully remote work arrangement.

β€’ Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers