
Technical Lead β GPU Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Romania.
β’ Take complete ownership of the Cosmic AC platform architecture from start to finish, which includes crafting architecture proposals and developing both high-level and low-level designs.
β’ Lead and manage a team of approximately twelve engineers distributed across backend, frontend, DevOps, QA, and documentation disciplines.
β’ Establish engineering standards, perform code and design reviews, manage release gates, facilitate one-on-one meetings, and provide input on growth and performance.
β’ Design, create, and maintain a managed Slurm service tailored for research users.
β’ Oversee controller and accounting functions, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drainage and autohealing, storage visibility, as well as identity and isolation.
β’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-supplied bare metal infrastructure.
β’ Implement and manage NVIDIA GPU Operator, Network Operator, VM-based GPU isolation with KubeVirt and VFIO, along with upgrades, backup, recovery, and node replacement processes.
β’ Own the architecture for managed inference, ensuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and capacity for confidential computing.
β’ Establish metrics, logging, alerting, and Service Level Objectives (SLOs) across the control plane, GPU fleet, and application tiers.
β’ Lead incident response initiatives, conduct post-incident reviews, and develop a sustainable on-call model.
β’ Act as the primary technical liaison to infrastructure partners and vendors.
β’ Transform requirements into detailed written specifications and acceptance tests, manage escalations to resolution, and contribute to capacity planning and hardware procurement.
β’ Translate research, model-training, and product workloads into platform requirements, while negotiating capacity when necessary.
β’ Complete the platform team and set a high technical standard for new engineers.
β’ A minimum of eight years of hands-on engineering experience.
β’ At least three years of experience in leading teams that build and manage infrastructure platforms.
β’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
β’ Practical experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and managing upgrades while jobs are running.
β’ Ideal candidates will have experience running HPC or GPU training clusters for research users.
β’ Familiarity with operating NVIDIA GPU fleets on bare metal, including managing driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
β’ Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance challenges.
β’ Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization techniques.
β’ Experience with production Kubernetes operations, covering control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy aspects.
β’ Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.
β’ Experience with tools like Prometheus, Grafana, Loki or their equivalents, as well as SLOs, incident response, and conducting post-incident reviews.
β’ Proficient in JavaScript and Node.js, with the ability to review and make architectural decisions concerning control-plane, CLI, and worker services.
β’ Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.
β’ Proven ability in people management across different time zones, conducting cross-track reviews, writing architectural decisions, and confidently challenging partners or executives with well-founded arguments.
β’ Excellent proficiency in written and spoken English.
β’ Ability to work from UTC through UTC+5:30 to facilitate collaboration with teams in Europe and India.
β’ Willingness to travel occasionally to partner sites and team events.
β’ Enjoy a fully remote work arrangement.
β’ Travel occasionally to partner sites and team events.
β’ Join a global, distributed team environment.
β’ Have the opportunity to work on cutting-edge digital finance, GPU infrastructure, AI, and blockchain platforms.
Verwaltungscloud.SH GmbH
WBS
EverCommerce
Get handpicked remote jobs straight to your inbox weekly.