HPC Infrastructure Engineer – GPU Clusters

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Manage and enhance the GPU fleet comprehensively, covering provisioning, scheduling, monitoring, upgrades, and capacity planning.

• Develop automation for node health monitoring, automated draining and remediation, as well as burn-in pipelines.

• Oversee the infrastructure stack supporting training code, which includes OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking.

• Execute and optimize job scheduling using Slurm or similar systems.

• Construct and sustain high-performance storage solutions for datasets and checkpoints.

• Diagnose and rectify performance issues related to stragglers, degraded links, thermal concerns, and malfunctioning GPUs.

• Assess rented GPU capacity by benchmarking providers, validating availability, and ensuring compliance with SLAs.

• Engage in hands-on hardware tasks, including racking, cabling, and diagnostics.

• Collaborate with datacenter personnel and vendors.

• Uphold cluster security through access control, network isolation, and secrets management.

• Partner directly with researchers to enhance training throughput and researcher efficiency.


⛳️ Requirements

• Experience in managing large-scale Linux server or GPU environments in a production setting.

• In-depth understanding of NVIDIA drivers, CUDA, NCCL, and DCGM, or extensive systems experience with the capacity to rapidly learn hardware stacks.

• Familiarity with bare-metal environments, server hardware, and high-speed networking.

• Proficient in Python and/or Bash scripting for automation purposes.

• Experience with infrastructure-as-code tools like Ansible or Terraform.

• Capability to analyze metrics, logs, and PromQL.

• Previous experience in supporting ML training workloads from an infrastructure perspective (preferred).

• Background in evaluating and collaborating with GPU cloud providers (preferred).

• Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (preferred).

• Familiarity with BMC/IPMI/Redfish automation and PXE provisioning at scale (preferred).

• Awareness of power and cooling requirements for dense GPU deployments (preferred).

• Willingness to engage in datacenter visits and perform hands-on hardware tasks.


🏝️ Benefits

• Annual discretionary professional development stipend.

• Annual discretionary social travel stipend for meeting colleagues.

• Annual company offsite event.

• Monthly co-working stipend for employees not located near a main hub.

• Flexibility for remote work.

• Opportunity to work from offices located in London, New York, San Francisco, or Warsaw.

People also viewed

Hitachi Solutions America1 day ago

Senior Azure Infrastructure Architect

DE flagGermany, +1 more countryFull-timeInfrastructure Engineer$140k – $180k/year
ApplyView job
Nava1 day ago

Senior Infrastructure Engineer, Azure

US flagAlabama, +29 more statesFull-timeInfrastructure Engineer$135.9k – $153k/year
ApplyView job
OpenTeams1 day ago

Senior Infrastructure Engineer – AI/ML Platform

US flagUnited States OnlyFull-timeInfrastructure Engineer$145k – $250k/year
ApplyView job
Kalepa1 day ago

Staff Platform Engineer – Infrastructure

US flagUnited States OnlyFull-timeInfrastructure Engineer$220k – $300k/year
ApplyView job
Activision Blizzard1 day ago

Senior Azure Infrastructure Engineer

US flagCalifornia OnlyFull-timeInfrastructure Engineer$102.8k – $190.2k/year
ApplyView job
Activision1 day ago

Senior Azure Infrastructure Engineer

US flagCalifornia OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers