
Technical Lead β GPU Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Ireland.
β’ Take ownership of the comprehensive architecture for Tether Data's Cosmic AC GPU compute and managed inference platform.
β’ Develop and uphold architecture proposals, high-level designs, and low-level designs through rigorous review processes.
β’ Lead and manage a team of approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation.
β’ Establish engineering standards, perform code and design reviews, oversee release gates, conduct one-on-ones, and provide feedback on growth and performance.
β’ Design, construct, and manage a Slurm service tailored for research users.
β’ Oversee controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, detection of stalled jobs and node health, draining, autohealing, storage visibility, identity, and isolation.
β’ Manage the bootstrap and lifecycle of the Kubernetes cluster on partner-supplied bare metal.
β’ Supervise the NVIDIA GPU Operator, Network Operator, VM-based GPU isolation, KubeVirt, VFIO, upgrades, backup and recovery, and node replacement.
β’ Lead the managed inference architecture, ensuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity.
β’ Implement metrics, logging, alerting, and SLOs across control plane, GPU fleet, and application tiers.
β’ Direct incident response efforts, conduct post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Convert requirements into detailed written specifications and acceptance tests, manage escalations, and assist in capacity planning and hardware sourcing.
β’ Collaborate with research, model-training, and product teams to translate workloads into platform requirements and negotiate limited capacity.
β’ Complete the platform team and set a high technical standard for new engineers.
β’ Take charge of implementation and delivery plans within a defined initial six-month delivery timeframe.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience in leading teams that build and maintain infrastructure platforms relied upon by other teams.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Demonstrated hands-on experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and performing upgrades while jobs are running.
β’ Ideally, experience managing an HPC or GPU training cluster for a research community.
β’ Proven experience operating NVIDIA GPU fleets on bare metal, encompassing driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
β’ Familiarity with InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance issues.
β’ In-depth knowledge of Linux systems, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization.
β’ Experience with production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design.
β’ Knowledge of HPC storage and data movement technologies, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Familiarity with Prometheus, Grafana, Loki or similar tools, as well as SLOs, incident response, and post-incident reviews.
β’ Proficiency in JavaScript and Node.js sufficient for reviewing control plane, CLI, and worker services and making architectural decisions.
β’ Experience delivering a multi-tenant IaaS, PaaS, or research computing service with resource isolation, quotas, usage metering, and user-friendly API and CLI interfaces.
β’ Proven ability in people management across time zones, cross-track reviews, written architectural decisions, and the capability to challenge partners or executives with well-founded reasoning.
β’ Excellent command of written and spoken English.
β’ Fully remote work location allowed, with a requirement to be situated between UTC and UTC+5:30.
β’ Willingness to undertake occasional travel to partner sites and team events.
β’ Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers.
β’ Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM.
β’ Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and relationships with hardware providers.
β’ Fully remote work environment.
β’ Opportunities for occasional travel to partner sites and team events.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.