
Senior HPC Cluster Administrator – Deep Learning Frameworks
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in Poland, +1 more state.
• Take charge of the entire lifecycle of GPU compute clusters, including procurement, provisioning, configuration management, monitoring, and decommissioning, within diverse Linux environments (DGX, HGX, embedded systems).
• Develop and expand storage solutions (NFS, Lustre, WekaFS, or similar) with a well-defined roadmap for capacity and performance enhancement.
• Spearhead the automation of infrastructure utilizing contemporary IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab).
• Oversee and optimize job scheduling through Slurm, encompassing fair-share policies, reservation management, and MIG/GPU partitioning strategies.
• Enhance and sustain observability stacks (Prometheus, Grafana, DCGM) while proactively addressing hardware and software incidents.
• Work in collaboration with ML engineers and software teams to fine-tune cluster configurations for large-scale distributed training tasks.
• Assess and integrate new technologies, including networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, and container runtimes, to boost performance and reliability.
• Mentor junior engineers and actively participate in establishing team-wide engineering standards.
• BS/MS in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
• Over 5 years of experience in deploying and managing large-scale HPC or ML training clusters.
• Extensive expertise in Linux systems administration at scale.
• Strong skills in scripting and automation using Python and/or Bash.
• Practical experience with Slurm, including scheduling, accounting, and cgroup configuration.
• Proficiency in configuration management and IaC, with Ansible required and Terraform as a plus.
• Familiarity with container technologies such as Docker, Apptainer/Singularity, and Kubernetes.
• Solid grasp of high-speed networking technologies (InfiniBand, RoCE, RDMA, EFA).
• Experience with distributed/parallel filesystems and storage architecture.
• Capability to manage problems from start to finish and communicate effectively with engineering and management stakeholders.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive team culture.
Moysies & Partner IT- und Managementberatung
The Cigna Group
WhiteWater Express Car Wash
Maximus
Get handpicked remote jobs straight to your inbox weekly.