Technical Lead – GPU Infrastructure

Posted 1 day ago

This is a fully remote position, open to applicants anywhere in the world.

📋 Description

• Take ownership of the complete platform architecture by creating architecture proposals, producing high-level and low-level designs, conducting reviews, and maintaining the baseline.

• Lead and manage a distributed team of approximately twelve engineers specializing in backend, frontend, DevOps, QA, and documentation.

• Define engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and provide input on growth and performance.

• Design, develop, and operate a managed Slurm service tailored for research users.

• Manage controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation.

• Oversee the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal.

• Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, along with upgrades, backup and recovery, and node replacement.

• Manage the architecture for inference, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

• Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application layers.

• Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.

• Act as the primary technical liaison to infrastructure partners and vendors.

• Convert requirements into documented specifications and acceptance tests, manage escalations, and assist with capacity planning and hardware procurement.

• Collaborate with research, model-training, and product teams to convert workloads into platform requirements and facilitate capacity brokering.

• Complete the platform team and set a high technical standard for new engineers.


⛳️ Requirements

• Eight or more years of hands-on engineering experience.

• A minimum of three years in leadership roles overseeing teams that develop and manage infrastructure platforms relied upon by other teams.

• Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

• Practical experience running Slurm at scale, which includes slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are present on the system.

• Experience in managing an HPC or GPU training cluster for a research community is highly preferred.

• Proficient in operating NVIDIA GPU fleets on bare metal, covering NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance.

• Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

• Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.

• Experience with production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.

• Knowledge of HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.

• Proficiency with Prometheus, Grafana, Loki or similar tools, including SLOs, incident response, and post-incident review processes.

• Working fluency in JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architectural decisions.

• Experience delivering a platform with real user engagement, such as a multi-tenant IaaS/PaaS or research computing service.

• People management experience across different time zones, cross-track review processes, written architectural decision-making, and communication with partners/executives.

• Excellent proficiency in written and spoken English.

• Must be located between UTC and UTC+5:30.

• Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

• Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.

• Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

• Fully remote work opportunity.

• Occasional travel to partner locations and team events.

People also viewed

CI&T18 hours ago

Senior Data Tech Lead

BR flagBrazil OnlyFull-timeFull-stack Engineer
ApplyView job
Prothera18 hours ago

Junior Full Stack Developer

BR flagBrazil OnlyFull-timeFull-stack Engineer
ApplyView job
The Home Depot18 hours ago

Senior Software Engineer

US flagUnited States OnlyFull-timeFull-stack Engineer$80k – $180k/year
ApplyView job
Truelogic Software18 hours ago

Senior Full Stack Engineer – Enterprise AI, Automation

DO flagDominican Republic OnlyFull-timeFull-stack Engineer
ApplyView job
Truelogic Software18 hours ago

Senior Full Stack Engineer – Enterprise AI & Automation

Latin AmericaFull-timeFull-stack Engineer
ApplyView job
Truelogic Software18 hours ago

Senior Full Stack Engineer – Enterprise AI, Automation

PA flagPanama OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers