
HPC Support Engineer
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in United States.
• Act as a senior technical escalation resource, diagnosing complex infrastructure and platform challenges down to the hardware, driver, or kernel level when necessary.
• Quickly and accurately differentiate between hardware failures, driver malfunctions, kernel-level issues, and misconfigurations in customer workloads to ensure issues are resolved accurately on the first attempt.
• Proactively identify and address gaps in processes, tools, and documentation, taking initiative to resolve them rather than waiting for assignments.
• Effectively utilize AI tools to develop scripts, automations, or small internal tools that bridge actual operational gaps (no professional development background needed).
• Conduct root-cause analysis across distributed systems, clusters, and GPU infrastructure.
• Create clear documentation for solutions and contribute to the enhancement of support procedures.
• Work collaboratively with engineering teams to transform recurring customer challenges into long-term solutions.
• Handle escalations from peers while providing training and mentorship throughout the process.
• Engage in a rotating on-call schedule, taking ownership of significant incidents and major customer issues.
• Be prepared to actively contribute wherever necessary, especially during rapid, high-volume deployments.
• Over 3 years of hands-on experience in HPC within an administration, support, or engineering capacity.
• Extensive understanding and experience in supporting Linux in a system administration role.
• Demonstrated experience in HPC environments, highlighting your expertise in Linux cluster administration, with a strong preference for Kubernetes and/or Slurm for managing clusters.
• Strong coding skills and experience with CI/CD, along with a proven record of utilizing AI-assisted tools for expedited workflows.
• Proficiency in monitoring and logging tools such as Prometheus, Grafana, and Datadog.
• Excellent skills in log analysis, debugging kernel-level problems, and performance profiling.
• Familiarity with CUDA, NCCL, NVLink, and GPUDirect RDMA.
• Experience with high-throughput networking technologies (IB/RoCE).
• Understanding of distributed AI/ML or HPC workloads.
• Knowledge of TCP/IP, VPNs, and firewalls within cloud environments.
• Capability to work independently and mentor junior support engineers.
• Health, dental, and vision insurance for you and your dependents.
• Wellness and commuter stipends available for select positions.
• 401k Plan with a 2% company match for employees in the USA.
• Flexible paid time off plan that is actively utilized by all staff.
Immersive Gamebox
Medlogix
alt.bank
Get handpicked remote jobs straight to your inbox weekly.