
HPC Infrastructure Engineer – GPU Clusters
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Manage and enhance the GPU fleet comprehensively, covering provisioning, scheduling, monitoring, upgrades, and capacity planning.
• Develop automation for node health monitoring, automated draining and remediation, as well as burn-in pipelines.
• Oversee the infrastructure stack supporting training code, which includes OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking.
• Execute and optimize job scheduling using Slurm or similar systems.
• Construct and sustain high-performance storage solutions for datasets and checkpoints.
• Diagnose and rectify performance issues related to stragglers, degraded links, thermal concerns, and malfunctioning GPUs.
• Assess rented GPU capacity by benchmarking providers, validating availability, and ensuring compliance with SLAs.
• Engage in hands-on hardware tasks, including racking, cabling, and diagnostics.
• Collaborate with datacenter personnel and vendors.
• Uphold cluster security through access control, network isolation, and secrets management.
• Partner directly with researchers to enhance training throughput and researcher efficiency.
• Experience in managing large-scale Linux server or GPU environments in a production setting.
• In-depth understanding of NVIDIA drivers, CUDA, NCCL, and DCGM, or extensive systems experience with the capacity to rapidly learn hardware stacks.
• Familiarity with bare-metal environments, server hardware, and high-speed networking.
• Proficient in Python and/or Bash scripting for automation purposes.
• Experience with infrastructure-as-code tools like Ansible or Terraform.
• Capability to analyze metrics, logs, and PromQL.
• Previous experience in supporting ML training workloads from an infrastructure perspective (preferred).
• Background in evaluating and collaborating with GPU cloud providers (preferred).
• Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (preferred).
• Familiarity with BMC/IPMI/Redfish automation and PXE provisioning at scale (preferred).
• Awareness of power and cooling requirements for dense GPU deployments (preferred).
• Willingness to engage in datacenter visits and perform hands-on hardware tasks.
• Annual discretionary professional development stipend.
• Annual discretionary social travel stipend for meeting colleagues.
• Annual company offsite event.
• Monthly co-working stipend for employees not located near a main hub.
• Flexibility for remote work.
• Opportunity to work from offices located in London, New York, San Francisco, or Warsaw.
Hitachi Solutions America
Nava
OpenTeams
Kalepa
Get handpicked remote jobs straight to your inbox weekly.