Technical Lead – GPU Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in Romania.

πŸ“‹ Description

β€’ Take complete ownership of the Cosmic AC platform architecture from start to finish, which includes crafting architecture proposals and developing both high-level and low-level designs.

β€’ Lead and manage a team of approximately twelve engineers distributed across backend, frontend, DevOps, QA, and documentation disciplines.

β€’ Establish engineering standards, perform code and design reviews, manage release gates, facilitate one-on-one meetings, and provide input on growth and performance.

β€’ Design, create, and maintain a managed Slurm service tailored for research users.

β€’ Oversee controller and accounting functions, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drainage and autohealing, storage visibility, as well as identity and isolation.

β€’ Manage the bootstrap and lifecycle of Kubernetes clusters on partner-supplied bare metal infrastructure.

β€’ Implement and manage NVIDIA GPU Operator, Network Operator, VM-based GPU isolation with KubeVirt and VFIO, along with upgrades, backup, recovery, and node replacement processes.

β€’ Own the architecture for managed inference, ensuring multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and capacity for confidential computing.

β€’ Establish metrics, logging, alerting, and Service Level Objectives (SLOs) across the control plane, GPU fleet, and application tiers.

β€’ Lead incident response initiatives, conduct post-incident reviews, and develop a sustainable on-call model.

β€’ Act as the primary technical liaison to infrastructure partners and vendors.

β€’ Transform requirements into detailed written specifications and acceptance tests, manage escalations to resolution, and contribute to capacity planning and hardware procurement.

β€’ Translate research, model-training, and product workloads into platform requirements, while negotiating capacity when necessary.

β€’ Complete the platform team and set a high technical standard for new engineers.


⛳️ Requirements

β€’ A minimum of eight years of hands-on engineering experience.

β€’ At least three years of experience in leading teams that build and manage infrastructure platforms.

β€’ A Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.

β€’ Practical experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and managing upgrades while jobs are running.

β€’ Ideal candidates will have experience running HPC or GPU training clusters for research users.

β€’ Familiarity with operating NVIDIA GPU fleets on bare metal, including managing driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.

β€’ Knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and troubleshooting multi-node NCCL performance challenges.

β€’ Extensive knowledge of Linux systems, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance optimization techniques.

β€’ Experience with production Kubernetes operations, covering control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy aspects.

β€’ Familiarity with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets.

β€’ Experience with tools like Prometheus, Grafana, Loki or their equivalents, as well as SLOs, incident response, and conducting post-incident reviews.

β€’ Proficient in JavaScript and Node.js, with the ability to review and make architectural decisions concerning control-plane, CLI, and worker services.

β€’ Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or research computing services.

β€’ Proven ability in people management across different time zones, conducting cross-track reviews, writing architectural decisions, and confidently challenging partners or executives with well-founded arguments.

β€’ Excellent proficiency in written and spoken English.

β€’ Ability to work from UTC through UTC+5:30 to facilitate collaboration with teams in Europe and India.

β€’ Willingness to travel occasionally to partner sites and team events.


🏝️ Benefits

β€’ Enjoy a fully remote work arrangement.

β€’ Travel occasionally to partner sites and team events.

β€’ Join a global, distributed team environment.

β€’ Have the opportunity to work on cutting-edge digital finance, GPU infrastructure, AI, and blockchain platforms.

People also viewed

Verwaltungscloud.SH GmbH19 hours ago

Senior Software Developer – Full-Stack

DE flagGermany OnlyFull-timeFull-stack Engineer€50k – €70k/year
ApplyView job
WBS19 hours ago

Linux/Application Administrator – Learning Platforms

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
ExactCare1 day ago

Senior Engineer

US flagOhio OnlyFull-timeFull-stack Engineer
ApplyView job
EverCommerce1 day ago

Senior Software Engineer – Growth

CA flagCanada, +1 more countryFull-timeFull-stack EngineerC$120k – C$150k/year
ApplyView job
PBS Radiology Business Experts1 day ago

Software Developer, C#/.NET

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Samsara1 day ago

Senior Software Engineer II – Tech Lead, External Platform

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers