
Staff HPC Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design, implement, and manage production Slurm clusters on both bare metal and virtual machines.
• Take responsibility for the architecture, lifecycle, multi-tenant scheduling policies, and reliability of GPU infrastructure within the Slurm cluster.
• Set up high availability, authentication, topology-aware GPU scheduling, accounting, partitions, QOS, fairshare, preemption, reservations, and tenant TRES limits.
• Lead the integration of the Slinky slurm-operator and assess the slurm-bridge for co-scheduling workloads on Kubernetes.
• Shift GPU nodes between Slurm batch-training queues and Kubernetes inference capabilities utilizing elastic-capacity mechanisms.
• Manage Pyxis/Enroot and OCI/containerd job paths, ensuring support for MPI/PMIx, module/Spack environments, and customer images.
• Develop both passive and active health checks for GPU clusters with automatic draining and job requeue functionalities.
• Oversee rack burn-in and acceptance testing prior to the scheduling of paid workloads.
• Provide reproducible clusters through the use of Terraform, Ansible, golden images, PXE, Redfish, and IPMI.
• Measure queue wait times, allocation efficiency, GPU utilization, and job failures using Prometheus/Grafana.
• Integrate Slurm accounting and GPU-hours with metering and invoicing workflows.
• Create runbooks and tenant documentation, onboard and support enterprise customers, manage escalations, and mentor platform engineers.
• 8+ years of experience in HPC, systems, or cloud infrastructure engineering.
• 4+ years managing production Slurm clusters with over 100 GPU nodes, serving real users and meeting service-level commitments.
• Extensive hands-on knowledge of slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions, QOS, fairshare, preemption, reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, as well as live-cluster version upgrades.
• Strong understanding of GPU and fabric fundamentals, including NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2, subnet manager/UFM, rail-optimized topology, GPUDirect RDMA, and NCCL tuning and failure diagnosis.
• Experience with production Kubernetes environments and familiarity with operator/CRD patterns.
• Hands-on experience with at least one Slurm-on-Kubernetes stack: Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator.
• Background in delivering bare-metal and virtualized compute solutions, including provisioning, firmware/BIOS lifecycle management, KVM/QEMU or public-cloud-equivalent VM clusters, and automation using Terraform/Ansible.
• Knowledge of parallel and shared storage solutions such as Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS.
• Proficient in Python and Bash for automating cluster tasks.
• Excellent written and verbal communication skills in English.
• Experience with Go is a plus.
• Equal employment opportunities in compliance with country, state, and local laws.
Anduril Industries
Sargent & Lundy
Sargent & Lundy
Get handpicked remote jobs straight to your inbox weekly.