Technical Lead – GPU Infrastructure

Posted 6 days ago

This is a fully remote position, open to applicants in Armenia.

📋 Description

• Take full ownership of the Cosmic AC platform architecture from start to finish, which includes creating architecture proposals, high-level and low-level designs, conducting reviews, and maintaining baselines.

• Lead and manage a geographically distributed team of around twelve engineers specializing in backend, frontend, DevOps, QA, and documentation.

• Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-one meetings, and provide insights on growth and performance.

• Design, build, and maintain a managed Slurm service tailored for research users.

• Oversee controller and accounting, manage partitions and login nodes, facilitate node onboarding, handle driver and CUDA baselines, and detect stalled jobs and node health issues, including drain and autohealing, along with ensuring storage visibility, identity, and isolation.

• Manage the bootstrap and lifecycle of Kubernetes clusters on bare metal provided by partners.

• Administer NVIDIA GPU Operator and Network Operator, including VM-based GPU isolation utilizing KubeVirt and VFIO, along with upgrades, backup, recovery, and node replacement.

• Control the architecture for managed inference, covering serving, multi-GPU and multi-node parallel processing, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

• Set up metrics, logging, alerting, and SLOs for the control plane, GPU fleet, and application tiers.

• Lead incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.

• Act as the primary technical liaison to infrastructure partners and vendors.

• Convert requirements into written specifications and acceptance tests, manage escalations, and assist with capacity planning and hardware sourcing.

• Collaborate with research, model-training, and product teams to transform workloads into platform requirements and broker capacity.

• Complete the platform team and establish the technical standards for onboarding new engineers.


⛳️ Requirements

• A minimum of eight years of hands-on engineering experience.

• At least three years of experience leading teams that build and manage infrastructure platforms relied upon by other teams.

• A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

• Practical experience with Slurm at scale, involving slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades while jobs are running.

• Preferred experience operating HPC or GPU training clusters for research users.

• Experience in managing NVIDIA GPU fleets on bare metal, including lifecycle management of drivers and CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.

• Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance challenges.

• Profound knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.

• Experience in production Kubernetes operations, encompassing control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.

• Background in HPC storage and data movement utilizing VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.

• Proficiency with Prometheus, Grafana, Loki or similar tools, SLOs, incident response, and conducting post-incident reviews.

• Working proficiency in JavaScript and Node.js sufficient to review and make architectural decisions regarding control-plane, CLI, and worker services.

• Prior experience in delivering a platform for real users, such as multi-tenant IaaS/PaaS or research computing services.

• Management of teams across different time zones, including cross-track reviews, written architecture decisions, and communication with partners and executives.

• Excellent written and spoken English communication skills.

• Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

• Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.

• Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, and experience with GPU-cloud/HPC/AI-lab environments, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

• Fully remote work arrangement.

• Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers