
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Egypt.
β’ Take ownership of the complete architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Lead the development of architecture proposals, high-level and low-level designs, conduct reviews, and maintain baselines.
β’ Manage and oversee approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation teams.
β’ Design, build, and maintain a managed Slurm service for research users.
β’ Oversee cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU and Network Operators, VM-based GPU isolation, and day-2 operations.
β’ Own the architecture for managed inference serving, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Establish metrics, logging, alerting, SLOs, incident response protocols, post-incident reviews, and a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors; define specifications, acceptance tests, escalations, capacity planning, and input for hardware sourcing.
β’ Convert research, model-training, and product workloads into platform requirements and manage capacity.
β’ Recruit and finalize the platform team while setting high technical standards.
β’ Ensure the platform is delivered within a specified six-month timeframe.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams responsible for building and maintaining infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Practical experience running Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority settings, accounting, prolog and epilog, node health scripting, and upgrades with active jobs.
β’ Ideally, experience operating an HPC or GPU training cluster for a research audience.
β’ Experience managing a bare-metal NVIDIA GPU fleet, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, and node burn-in and acceptance processes.
β’ Knowledge of InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Expertise in Linux systems, covering kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience with production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Familiarity with HPC storage and data movement utilizing shared filesystems such as VAST, Lustre, and NFS, node-local NVMe caching, and the distribution of large model weights and datasets.
β’ Proficiency in observability and operations tools like Prometheus, Grafana, Loki or similar, along with SLOs, incident response, and post-incident review processes.
β’ Working proficiency in JavaScript and Node.js for evaluating and making architectural decisions regarding control plane, CLI, and worker services.
β’ Proven experience delivering a platform used by real users, such as multi-tenant IaaS/PaaS or research computing services.
β’ Experience in people management across different time zones, conducting cross-track reviews, documenting architecture decisions, and the ability to communicate technical decisions to partners and executives.
β’ Excellent proficiency in written and spoken English.
β’ Desirable experience includes knowledge of Slurm operators, Kubernetes-native schedulers, modern serving stacks, VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and partnerships with hardware providers.
β’ Remote work opportunities.
β’ Occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.