Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in Egypt.

πŸ“‹ Description

β€’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.

β€’ Lead the development of architecture proposals, high-level and low-level designs, conduct reviews, and maintain baselines.

β€’ Manage and oversee approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation teams.

β€’ Design, build, and maintain a managed Slurm service for research users.

β€’ Oversee cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU and Network Operators, VM-based GPU isolation, and day-2 operations.

β€’ Own the architecture for managed inference serving, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.

β€’ Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and a sustainable on-call model.

β€’ Act as the primary technical liaison to infrastructure partners and vendors; define specifications, acceptance tests, escalations, capacity planning, and input for hardware sourcing.

β€’ Convert research, model-training, and product workloads into platform requirements and manage capacity.

β€’ Recruit and finalize the platform team while setting high technical standards.

β€’ Ensure the platform is delivered within a specified six-month timeframe.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience.

β€’ At least three years of experience leading teams responsible for building and maintaining infrastructure platforms relied upon by other teams.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Practical experience running Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority settings, accounting, prolog and epilog, node health scripting, and upgrades with active jobs.

β€’ Ideally, experience operating an HPC or GPU training cluster for a research audience.

β€’ Experience managing a bare-metal NVIDIA GPU fleet, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, and node burn-in and acceptance processes.

β€’ Knowledge of InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

β€’ Expertise in Linux systems, covering kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.

β€’ Experience with production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.

β€’ Familiarity with HPC storage and data movement utilizing shared filesystems such as VAST, Lustre, and NFS, node-local NVMe caching, and the distribution of large model weights and datasets.

β€’ Proficiency in observability and operations tools like Prometheus, Grafana, Loki or similar, along with SLOs, incident response, and post-incident review processes.

β€’ Working proficiency in JavaScript and Node.js for evaluating and making architectural decisions regarding control plane, CLI, and worker services.

β€’ Proven experience delivering a platform used by real users, such as multi-tenant IaaS/PaaS or research computing services.

β€’ Experience in people management across different time zones, conducting cross-track reviews, documenting architecture decisions, and the ability to communicate technical decisions to partners and executives.

β€’ Excellent proficiency in written and spoken English.

β€’ Desirable experience includes knowledge of Slurm operators, Kubernetes-native schedulers, modern serving stacks, VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

β€’ Remote work opportunities.

β€’ Occasional travel to partner sites and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers