
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Poland.
β’ Take ownership of the complete platform architecture through proposals, high-level and low-level designs, reviews, and maintaining the baseline.
β’ Lead and manage a distributed team of around twelve engineers specializing in backend, frontend, DevOps, QA, and documentation.
β’ Set engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide input on growth and performance.
β’ Design, construct, and maintain a managed Slurm service for research users.
β’ Oversee Slurm controllers, accounting, partitions, login nodes, onboarding and acceptance of nodes, driver and CUDA baselines and upgrades, stalled-job and node-health detection, as well as drain and autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on bare metal provided by partners.
β’ Implement NVIDIA GPU Operator and Network Operator, establish VM-based GPU isolation using KubeVirt and VFIO, manage upgrades, backup and recovery, and node replacement.
β’ Take charge of managed inference architecture, which includes serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into detailed written specifications and acceptance tests, manage escalations to resolution, and contribute to capacity planning and hardware procurement.
β’ Collaborate with research, model-training, and product teams to convert workloads into platform specifications and facilitate capacity brokering.
β’ Complete the platform team while establishing the technical expectations for new engineers.
β’ Ensure the delivery of the stack within a defined timeframe during the initial six months.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams responsible for building and operating infrastructure platforms that other teams rely on.
β’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Proven hands-on experience with Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.
β’ Ideally, experience in operating an HPC or GPU training cluster for a research community.
β’ Experience with managing NVIDIA GPU fleets on bare metal, which includes NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in and acceptance.
β’ Familiarity with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Comprehensive understanding of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience in production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Knowledge of HPC storage and data movement technologies, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Proficiency with tools like Prometheus, Grafana, Loki or equivalents, along with experience in SLOs, incident response, and post-incident reviews.
β’ Working proficiency in JavaScript and Node.js sufficient to evaluate control plane, CLI, and worker services and make architectural decisions.
β’ Experience delivering a multi-tenant IaaS, PaaS, or research computing service that includes resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
β’ Effective people management skills across various time zones, ability to conduct cross-track reviews, document architectural decisions, and challenge partners or executives when necessary.
β’ Excellent command of written and spoken English.
β’ Availability to work within the UTC to UTC+5:30 time frame.
β’ Preferred experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Preferred experience with contemporary serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Preferred experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships.
β’ Fully remote work arrangement.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.