Technical Lead – GPU Infrastructure

Posted 6 days ago

This is a fully remote position, open to applicants in Egypt.

πŸ“‹ Description

β€’ Take full responsibility for the architecture of the Cosmic AC GPU compute and managed inference platform from start to finish.

β€’ Create and sustain architecture proposals, high-level designs, and low-level designs through thorough reviews.

β€’ Oversee and manage a team of around twelve distributed engineers across various domains, including backend, frontend, DevOps, QA, and documentation.

β€’ Set engineering standards, perform code and design reviews, manage release gates, conduct one-on-one meetings, and provide feedback for growth and performance.

β€’ Design, develop, and manage a Slurm service tailored for research users.

β€’ Handle the Slurm controller, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health monitoring, autohealing, storage visibility, identity, and isolation.

β€’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-supplied bare metal.

β€’ Operate NVIDIA GPU Operator and Network Operator, KubeVirt, and implement VFIO-based GPU isolation, upgrades, backup, recovery, and node replacement.

β€’ Architect managed inference at scale, covering serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential compute-capable capacity.

β€’ Define metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers.

β€’ Lead incident response efforts, conduct post-incident reviews, and implement a sustainable on-call model.

β€’ Serve as the primary technical liaison to infrastructure partners and vendors.

β€’ Convert requirements into written specifications and acceptance tests, handle escalations, and assist in capacity planning and hardware sourcing.

β€’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and negotiate capacity.

β€’ Complete the platform team and establish the technical standards for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience.

β€’ At least three years of experience leading teams responsible for building and maintaining infrastructure platforms relied upon by other teams.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Direct experience operating Slurm at scale, including components such as slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, health scripting, and upgrades while jobs are running.

β€’ Ideal candidates will have experience managing an HPC or GPU training cluster for a research community.

β€’ Experience with operating NVIDIA GPU fleets on bare metal, including managing the driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance testing.

β€’ Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.

β€’ Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.

β€’ Proficient in production Kubernetes operations, including control plane management, upgrades, CNI, CSI, operators, custom controllers, and designing for multi-tenancy.

β€’ Experience with HPC storage and data transfer, including technologies such as VAST, Lustre, NFS, node-local NVMe caching, and the distribution of large model weights and datasets.

β€’ Familiarity with monitoring tools such as Prometheus, Grafana, Loki or similar, as well as SLOs, incident response, and post-incident analysis.

β€’ Working knowledge of JavaScript and Node.js sufficient for reviewing control-plane, CLI, and worker services and making architecture decisions.

β€’ Experience deploying a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.

β€’ Competence in managing teams across different time zones, conducting cross-track reviews, documenting architecture decisions, and effectively challenging partners or executives with rationale.

β€’ Excellent written and spoken English communication skills.

β€’ Must be available to work within the UTC to UTC+5:30 time zones.

β€’ Desirable: Experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.

β€’ Desirable: Familiarity with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.

β€’ Desirable: Knowledge of VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and partnerships with hardware providers.


🏝️ Benefits

β€’ Fully remote work arrangement.

β€’ Occasional travel to partner locations and team events.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers