Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in Armenia.

πŸ“‹ Description

β€’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.

β€’ Drive the development of architecture proposals, along with high-level and low-level designs, reviews, and the upkeep of the architectural baseline.

β€’ Lead and manage a distributed team of approximately twelve engineers across backend, frontend, DevOps, QA, and documentation.

β€’ Establish engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide feedback on growth and performance.

β€’ Design, develop, and maintain a managed Slurm service for research users.

β€’ Oversee the Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drainage and autohealing, storage visibility, identity, and isolation.

β€’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations.

β€’ Design the managed inference serving architecture, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and capabilities for confidential computing.

β€’ Set up metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.

β€’ Engage in incident response, conduct post-incident reviews, and develop a sustainable on-call model.

β€’ Act as the primary technical liaison to infrastructure partners and vendors.

β€’ Convert requirements into written specifications and acceptance tests, manage escalations to resolution, and contribute to capacity planning and hardware sourcing.

β€’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and facilitate capacity provisioning.

β€’ Complete the platform team and set high technical standards for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience, with at least three years in leadership roles for teams that build and maintain infrastructure platforms relied upon by other teams.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Proven hands-on experience operating Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority settings, accounting, prolog and epilog, node health scripting, and performing upgrades with active jobs in the system.

β€’ Ideally, experience operating an HPC or GPU training cluster for a research audience is preferred.

β€’ Experience managing NVIDIA GPU fleets on bare metal, including lifecycle management of NVIDIA drivers and CUDA, Fabric Manager and NVSwitch behavior, DCGM-based health and utilization metrics, MIG, and node burn-in and acceptance testing.

β€’ Familiarity with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

β€’ In-depth knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, as well as performance tuning.

β€’ Experience with production Kubernetes operation, including control plane management, upgrades, CNI and CSI, custom operators and controllers, and multi-tenancy strategies.

β€’ Knowledge of HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.

β€’ Experience with monitoring tools like Prometheus, Grafana, and Loki or their equivalents, as well as SLOs, incident response, and post-incident reviews.

β€’ Proficient in JavaScript and Node.js, capable of reviewing and making architectural decisions regarding control plane, CLI, and worker services.

β€’ Experience delivering a deployed platform with actual users, such as a multi-tenant IaaS/PaaS or research computing service, encompassing resource isolation, quotas, usage tracking, and user-facing API and CLI interfaces.

β€’ Experience in managing people across different time zones, conducting cross-track reviews, documenting architectural decisions, and effectively communicating those decisions to partners or executives.

β€’ Excellent proficiency in both written and spoken English.

β€’ Must be based within UTC and UTC+5:30 time zones.

β€’ Preferred experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

β€’ Preferred familiarity with vLLM, SGLang, TensorRT-LLM, GPU isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts with acceptance tests.


🏝️ Benefits

β€’ Fully remote work opportunity.

β€’ Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers