
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Egypt.
β’ Take full responsibility for the architecture of the Cosmic AC GPU compute and managed inference platform from start to finish.
β’ Create and sustain architecture proposals, high-level designs, and low-level designs through thorough reviews.
β’ Oversee and manage a team of around twelve distributed engineers across various domains, including backend, frontend, DevOps, QA, and documentation.
β’ Set engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide feedback for growth and performance.
β’ Design, develop, and manage a Slurm service tailored for research users.
β’ Handle the Slurm controller, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health monitoring, autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-supplied bare metal.
β’ Operate NVIDIA GPU Operator and Network Operator, KubeVirt, and implement VFIO-based GPU isolation, upgrades, backup, recovery, and node replacement.
β’ Architect managed inference at scale, covering serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential compute-capable capacity.
β’ Define metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.
β’ Lead incident response efforts, conduct post-incident reviews, and implement a sustainable on-call model.
β’ Serve as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into written specifications and acceptance tests, handle escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and negotiate capacity.
β’ Complete the platform team and establish the technical standards for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience leading teams responsible for building and maintaining infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Direct experience operating Slurm at scale, including components such as slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, health scripting, and upgrades while jobs are running.
β’ Ideal candidates will have experience managing an HPC or GPU training cluster for a research community.
β’ Experience with operating NVIDIA GPU fleets on bare metal, including managing the driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance testing.
β’ Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Proficient in production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and designing for multi-tenancy.
β’ Experience with HPC storage and data transfer, including technologies such as VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.
β’ Familiarity with monitoring tools such as Prometheus, Grafana, Loki or similar, as well as SLOs, incident response, and post-incident analysis.
β’ Working knowledge of JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architecture decisions.
β’ Experience deploying a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.
β’ Competence in managing teams across different time zones, conducting cross-track reviews, documenting architecture decisions, and effectively challenging partners or executives with rationale.
β’ Excellent written and spoken English communication skills.
β’ Must be available to work within the UTC to UTC+5:30 time zones.
β’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.
β’ Fully remote work arrangement.
β’ Occasional travel to partner locations and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.