
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
β’ Take full ownership of the Cosmic AC platform architecture from beginning to end by creating architecture proposals, high-level and low-level designs, conducting reviews, and maintaining baselines.
β’ Manage and lead approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Set engineering standards, perform code and design reviews, oversee release gates, conduct one-on-one meetings, and provide input on growth and performance.
β’ Design, develop, and manage a Slurm service tailored for research users.
β’ Oversee Slurm controller/accounting, partitions, login nodes, node onboarding, driver/CUDA baselines, detection of stalled jobs and node health, drainage, autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations.
β’ Define the architecture for managed inference, addressing multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and capacity for confidential computing.
β’ Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and sustainable on-call practices.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations to resolution, and assist with capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and facilitate capacity brokering.
β’ Complete the platform team and establish technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams that build and maintain infrastructure platforms relied upon by other teams.
β’ Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority settings, accounting, prolog and epilog, node health scripting, and managing upgrades with jobs actively running on the system.
β’ Ideally, experience operating an HPC or GPU training cluster for a research community is required.
β’ Proven experience managing NVIDIA GPU fleets on bare metal, encompassing NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
β’ Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and resolving multi-node NCCL performance issues.
β’ In-depth knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience in production Kubernetes operations, covering control plane management, upgrades, CNI, CSI, operators, custom controllers, and designing for multi-tenancy.
β’ Background in HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Experience with monitoring tools such as Prometheus, Grafana, Loki, or equivalents, along with SLOs, incident response, and post-incident reviews.
β’ Proficiency in JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architectural decisions.
β’ Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.
β’ Demonstrated ability in people management across time zones, conducting cross-track reviews, documenting architectural decisions, and exercising technical judgment in collaboration with partners and executives.
β’ Excellent written and spoken English skills.
β’ Fully remote role, available for candidates located between UTC and UTC+5:30.
β’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable: Knowledge of VM/container isolation and confidential computing technologies.
β’ Desirable: Experience with Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, and GitOps practices.
β’ Desirable: Background in GPU cloud, HPC centers, AI lab platforms, peer-to-peer, or distributed systems.
β’ Desirable: Experience in managing hardware-provider relationships and handling written contracts/acceptance tests.
β’ Fully remote work arrangement.
β’ Opportunities for occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.