
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Italy.
β’ Take responsibility for the complete platform architecture through proposals, high-level and low-level design, reviews, and upkeep of the baseline.
β’ Lead and manage a distributed team of around twelve engineers across backend, frontend, DevOps, QA, and documentation disciplines.
β’ Establish engineering standards, conduct code and design reviews, oversee release gates, hold one-on-one meetings, and provide insights on growth and performance.
β’ Design, construct, and maintain a managed Slurm service for research users.
β’ Oversee Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of the Kubernetes cluster on partner-provided bare metal.
β’ Handle the NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement.
β’ Own the managed inference architecture, encompassing serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Set up metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, conduct post-incident reviews, and design a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations, and provide advice on capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and broker capacity.
β’ Complete the platform team and set the technical standards for new engineers.
β’ Manage architecture, implementation, and delivery plans within a fixed initial six-month delivery timeframe.
β’ Eight or more years of practical engineering experience.
β’ A minimum of three years leading teams that develop and operate infrastructure platforms relied upon by other teams.
β’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Hands-on experience managing Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and performing upgrades while jobs are running.
β’ Ideally, experience operating an HPC or GPU training cluster for a research audience.
β’ Experience with bare-metal NVIDIA GPU fleet management, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance testing.
β’ Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Strong understanding of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience in production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.
β’ Familiarity with HPC storage and data movement using VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Observability and operations experience with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and post-incident reviews.
β’ Proficiency in JavaScript and Node.js sufficient to evaluate control-plane, CLI, and worker services and make architectural decisions.
β’ Experience delivering a multi-tenant IaaS, PaaS, or research computing service with isolation, quotas, usage metering, and user-facing API and CLI interfaces.
β’ Management experience across time zones, conducting cross-track reviews, writing architectural decisions, and the ability to challenge partners or executives with sound reasoning.
β’ Excellent written and spoken English communication skills.
β’ Must reside between UTC and UTC+5:30.
β’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable: Knowledge of VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, and experience with GPU cloud/HPC/AI lab environments, distributed systems, and partnerships with hardware providers.
β’ Fully remote work arrangement.
β’ Occasional travel to partner locations and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.