
Technical Lead – GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Armenia.
• Take full ownership of the Cosmic AC platform architecture from start to finish, which includes creating architecture proposals, high-level and low-level designs, conducting reviews, and maintaining baselines.
• Lead and manage a geographically distributed team of around twelve engineers specializing in backend, frontend, DevOps, QA, and documentation.
• Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-one meetings, and provide insights on growth and performance.
• Design, build, and maintain a managed Slurm service tailored for research users.
• Oversee controller and accounting, manage partitions and login nodes, facilitate node onboarding, handle driver and CUDA baselines, and detect stalled jobs and node health issues, including drain and autohealing, along with ensuring storage visibility, identity, and isolation.
• Manage the bootstrap and lifecycle of Kubernetes clusters on bare metal provided by partners.
• Administer NVIDIA GPU Operator and Network Operator, including VM-based GPU isolation utilizing KubeVirt and VFIO, along with upgrades, backup, recovery, and node replacement.
• Control the architecture for managed inference, covering serving, multi-GPU and multi-node parallel processing, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
• Set up metrics, logging, alerting, and SLOs for the control plane, GPU fleet, and application tiers.
• Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.
• Act as the primary technical liaison to infrastructure partners and vendors.
• Convert requirements into written specifications and acceptance tests, manage escalations, and assist with capacity planning and hardware sourcing.
• Collaborate with research, model-training, and product teams to transform workloads into platform requirements and broker capacity.
• Complete the platform team and establish the technical standards for onboarding new engineers.
• A minimum of eight years of hands-on engineering experience.
• At least three years of experience leading teams that build and manage infrastructure platforms relied upon by other teams.
• A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
• Practical experience with Slurm at scale, involving slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.
• Preferred experience operating HPC or GPU training clusters for research users.
• Experience in managing NVIDIA GPU fleets on bare metal, including lifecycle management of drivers and CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
• Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance challenges.
• Profound knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
• Experience in production Kubernetes operations, encompassing control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.
• Background in HPC storage and data movement utilizing VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
• Proficiency with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and conducting post-incident reviews.
• Working proficiency in JavaScript and Node.js sufficient to review and make architectural decisions regarding control-plane, CLI, and worker services.
• Prior experience in delivering a platform for real users, such as multi-tenant IaaS/PaaS or research computing services.
• Management of teams across different time zones, including cross-track reviews, written architecture decisions, and communication with partners and executives.
• Excellent written and spoken English communication skills.
• Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
• Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
• Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, and experience with GPU-cloud/HPC/AI-lab environments, distributed systems, and partnerships with hardware providers.
• Fully remote work arrangement.
• Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.