
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Armenia.
β’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Drive the development of architecture proposals, along with high-level and low-level designs, reviews, and the upkeep of the architectural baseline.
β’ Lead and manage a distributed team of approximately twelve engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide feedback on growth and performance.
β’ Design, develop, and maintain a managed Slurm service for research users.
β’ Oversee the Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drainage and autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations.
β’ Design the managed inference serving architecture, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and capabilities for confidential computing.
β’ Set up metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Engage in incident response, conduct post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, manage escalations to resolution, and contribute to capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and facilitate capacity provisioning.
β’ Complete the platform team and set high technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience, with at least three years in leadership roles for teams that build and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Proven hands-on experience operating Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority settings, accounting, prolog and epilog, node health scripting, and performing upgrades with active jobs in the system.
β’ Ideally, experience operating an HPC or GPU training cluster for a research audience is preferred.
β’ Experience managing NVIDIA GPU fleets on bare metal, including lifecycle management of NVIDIA drivers and CUDA, Fabric Manager and NVSwitch behavior, DCGM-based health and utilization metrics, MIG, and node burn-in and acceptance testing.
β’ Familiarity with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ In-depth knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, as well as performance tuning.
β’ Experience with production Kubernetes operation, including control plane management, upgrades, CNI and CSI, custom operators and controllers, and multi-tenancy strategies.
β’ Knowledge of HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.
β’ Experience with monitoring tools like Prometheus, Grafana, and Loki or their equivalents, as well as SLOs, incident response, and post-incident reviews.
β’ Proficient in JavaScript and Node.js, capable of reviewing and making architectural decisions regarding control plane, CLI, and worker services.
β’ Experience delivering a deployed platform with actual users, such as a multi-tenant IaaS/PaaS or research computing service, encompassing resource isolation, quotas, usage tracking, and user-facing API and CLI interfaces.
β’ Experience in managing people across different time zones, conducting cross-track reviews, documenting architectural decisions, and effectively communicating those decisions to partners or executives.
β’ Excellent proficiency in both written and spoken English.
β’ Must be based within UTC and UTC+5:30 time zones.
β’ Preferred experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Preferred familiarity with vLLM, SGLang, TensorRT-LLM, GPU isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts with acceptance tests.
β’ Fully remote work opportunity.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.