
Technical Support Engineer β Slurm
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in India.
β’ Take ownership of Slurm support cases from the initial investigation to resolution for clients operating production AI and HPC clusters.
β’ Troubleshoot intricate issues related to slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and ensuring high availability.
β’ Address Slurm configuration and policy challenges involving partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
β’ Analyze performance, reliability, and scalability concerns utilizing logs, diagnostic data, configuration assessments, reproduction efforts, and source-level debugging as necessary.
β’ Identify issues across Slurm and its dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster management systems.
β’ Provide guidance to customers on Slurm configuration, upgrades, operational practices, resource management, and safe recovery procedures from production incidents.
β’ Work collaboratively with engineering teams by providing clear technical descriptions, reproducible test cases, and well-documented defect reports.
β’ Create guides, knowledge base articles, diagnostic tools, and internal training materials to enhance Slurm expertise throughout the support organization.
β’ Bachelorβs degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
β’ More than 5 years of practical experience administering and supporting Slurm in production HPC or AI settings, including handling business-critical outage incidents.
β’ Advanced understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
β’ Capability to independently diagnose complex Slurm incidents and lead them to sound technical resolutions.
β’ Extensive experience in Linux system administration and troubleshooting, including systemd, cgroups, authentication, networking, and database-backed services.
β’ Experience managing Slurm in multi-user clusters with intricate scheduling policies and diverse computing resources.
β’ Strong analytical and research skills, especially in differentiating Slurm defects from configuration, integration, infrastructure, and workload challenges.
β’ Exceptional written and verbal communication abilities, particularly in transforming detailed technical findings into clear explanations and actionable recommendations.
β’ Experience in supporting large-scale Slurm environments that consist of thousands of nodes or GPUs.
β’ Proficiency in diagnosing scheduler performance, job throughput, controller load, and database scaling issues.
β’ Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.
β’ Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.
β’ Prior experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or any other cluster management platform.
β’ A diverse and supportive work environment.
β’ Opportunities to create a lasting impact on the world.
β’ Full-time employment.
US CLOUD: Microsoft Premier/Unified Support Alternative
Intuitive
Bunzl Distribution NA
Palo Alto Networks
Get handpicked remote jobs straight to your inbox weekly.