Remotery

HPC Support Engineer

atLambdaRemoteUS flagUnited StatesFull-timeSupport EngineerMid-levelSenior$122k – $162k/year

Posted Jul 31

This is a fully remote position, open to applicants in United States.

📋 Description

• Act as a senior technical escalation resource, diagnosing complex infrastructure and platform challenges down to the hardware, driver, or kernel level when necessary.

• Quickly and accurately differentiate between hardware failures, driver malfunctions, kernel-level issues, and misconfigurations in customer workloads to ensure issues are resolved accurately on the first attempt.

• Proactively identify and address gaps in processes, tools, and documentation, taking initiative to resolve them rather than waiting for assignments.

• Effectively utilize AI tools to develop scripts, automations, or small internal tools that bridge actual operational gaps (no professional development background needed).

• Conduct root-cause analysis across distributed systems, clusters, and GPU infrastructure.

• Create clear documentation for solutions and contribute to the enhancement of support procedures.

• Work collaboratively with engineering teams to transform recurring customer challenges into long-term solutions.

• Handle escalations from peers while providing training and mentorship throughout the process.

• Engage in a rotating on-call schedule, taking ownership of significant incidents and major customer issues.

• Be prepared to actively contribute wherever necessary, especially during rapid, high-volume deployments.


⛳️ Requirements

• Over 3 years of hands-on experience in HPC within an administration, support, or engineering capacity.

• Extensive understanding and experience in supporting Linux in a system administration role.

• Demonstrated experience in HPC environments, highlighting your expertise in Linux cluster administration, with a strong preference for Kubernetes and/or Slurm for managing clusters.

• Strong coding skills and experience with CI/CD, along with a proven record of utilizing AI-assisted tools for expedited workflows.

• Proficiency in monitoring and logging tools such as Prometheus, Grafana, and Datadog.

• Excellent skills in log analysis, debugging kernel-level problems, and performance profiling.

• Familiarity with CUDA, NCCL, NVLink, and GPUDirect RDMA.

• Experience with high-throughput networking technologies (IB/RoCE).

• Understanding of distributed AI/ML or HPC workloads.

• Knowledge of TCP/IP, VPNs, and firewalls within cloud environments.

• Capability to work independently and mentor junior support engineers.


🏝️ Benefits

• Health, dental, and vision insurance for you and your dependents.

• Wellness and commuter stipends available for select positions.

• 401k Plan with a 2% company match for employees in the USA.

• Flexible paid time off plan that is actively utilized by all staff.

People also viewed

Immersive Gamebox4 days ago

Technical Support, Guest Experience Specialist

GB flagUnited Kingdom OnlyFull-timeSupport Engineer£30k/year
ApplyView job
Medlogix4 days ago

Application Support Analyst – MST/PST Time Zone

US flagUnited States OnlyFull-timeSupport Engineer$0 – $85k/year
ApplyView job
alt.bank4 days ago

Estagiário em Suporte Técnico N3

BR flagBrazil OnlyInternshipSupport Engineer
ApplyView job
OpenAI4 days ago

Senior Support Engineer

CA flagCanada OnlyFull-timeSupport Engineer
ApplyView job
OpenAI4 days ago

AI Support Engineer

CA flagCanada OnlyFull-timeSupport Engineer
ApplyView job
Aurum Software4 days ago

Analista de Suporte Júnior

BR flagBrazil OnlyFull-timeSupport Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers