
Senior Technical Support Engineer – Slurm
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in Texas.
• Take ownership of Slurm support cases from the initial investigation through to resolution for clients operating production AI and HPC clusters.
• Analyze intricate issues related to slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.
• Resolve Slurm configuration and policy challenges concerning partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
• Examine performance, reliability, and scalability challenges using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging as needed.
• Identify problems across Slurm and its dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.
• Provide guidance to customers on Slurm configuration, upgrades, operational practices, system-resource management, and safe recovery from production incidents.
• Work collaboratively with engineering teams by creating technical descriptions, reproducible test cases, and defect reports.
• Author guides, knowledge-base articles, diagnostic tools, and internal training materials to enhance Slurm expertise within the support organization.
• Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent professional experience.
• More than 5 years of practical experience administering and supporting Slurm in production HPC or AI settings, including handling business-critical outage incidents.
• Comprehensive understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
• Ability to independently identify complex Slurm incidents and guide them to a technically sound resolution.
• Extensive experience in Linux system administration and troubleshooting, including systemd, cgroups, authentication, networking, and database-backed services.
• Proven experience managing Slurm across multi-user clusters with intricate scheduling policies and diverse compute resources.
• Strong analytical and research capabilities, including the ability to differentiate Slurm defects from configuration, integration, infrastructure, and workload issues.
• Exceptional written and verbal communication abilities.
• Experience supporting large-scale Slurm environments comprising thousands of nodes or GPUs.
• Proficiency in diagnosing scheduler performance, job throughput, controller load, and database scaling challenges.
• Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.
• Experience with container technologies and HPC integration tools such as Pyxis, Enroot, Apptainer, or Singularity.
• Past experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster management platform.
• Highly competitive salaries.
• Comprehensive benefits package.
• Equity.
Contajá Contabilidade Online
Contajá Contabilidade Online
Vultr
DaVita Kidney Care
Get handpicked remote jobs straight to your inbox weekly.