Remotery

Senior HPC Cluster Administrator – Deep Learning Frameworks

atNVIDIARemoteFull-timeAdministrationSeniorPLN 221.3k – PLN 507k/year

Posted Jul 22

This is a fully remote position, open to applicants in Poland, +1 more state.

📋 Description

• Take charge of the entire lifecycle of GPU compute clusters, including procurement, provisioning, configuration management, monitoring, and decommissioning, within diverse Linux environments (DGX, HGX, embedded systems).

• Develop and expand storage solutions (NFS, Lustre, WekaFS, or similar) with a well-defined roadmap for capacity and performance enhancement.

• Spearhead the automation of infrastructure utilizing contemporary IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab).

• Oversee and optimize job scheduling through Slurm, encompassing fair-share policies, reservation management, and MIG/GPU partitioning strategies.

• Enhance and sustain observability stacks (Prometheus, Grafana, DCGM) while proactively addressing hardware and software incidents.

• Work in collaboration with ML engineers and software teams to fine-tune cluster configurations for large-scale distributed training tasks.

• Assess and integrate new technologies, including networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, and container runtimes, to boost performance and reliability.

• Mentor junior engineers and actively participate in establishing team-wide engineering standards.


⛳️ Requirements

• BS/MS in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.

• Over 5 years of experience in deploying and managing large-scale HPC or ML training clusters.

• Extensive expertise in Linux systems administration at scale.

• Strong skills in scripting and automation using Python and/or Bash.

• Practical experience with Slurm, including scheduling, accounting, and cgroup configuration.

• Proficiency in configuration management and IaC, with Ansible required and Terraform as a plus.

• Familiarity with container technologies such as Docker, Apptainer/Singularity, and Kubernetes.

• Solid grasp of high-speed networking technologies (InfiniBand, RoCE, RDMA, EFA).

• Experience with distributed/parallel filesystems and storage architecture.

• Capability to manage problems from start to finish and communicate effectively with engineering and management stakeholders.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and continuous learning.

• Collaborative and inclusive team culture.

People also viewed

Moysies & Partner IT- und Managementberatung1 day ago

Director – Business Center Digital State & Public Administration

DE flagGermany OnlyFull-timeAdministration
ApplyView job
The Cigna Group2 days ago

340B Implementation Administrator

US flagUnited States OnlyFull-timeAdministration$87.8k – $146.4k/year
ApplyView job
WhiteWater Express Car Wash2 days ago

Point of Sale (POS) Administrator

US flagUnited States OnlyFull-timeAdministration$60k – $70k/year
ApplyView job
Maximus2 days ago

Clinical Administrator

US flagUnited States OnlyFull-timeAdministration$18 – $20/hour
ApplyView job
EXL2 days ago

Assistant Manager, Administration and Projects – Healthcare Payment Integrity

US flagUnited States OnlyFull-timeAdministration$23 – $38/hour
ApplyView job
FCamara Consulting & Training2 days ago

Administrador de bases de dados, DBA

BR flagBrazil OnlyFull-timeAdministration
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers