Remotery

Technical Support Engineer – GPU Clusters

atTogether AIRemoteUS flagCaliforniaFull-timeSupport EngineerMid-levelSenior$160k – $230k/year

Posted Aug 5

This is a fully remote position, open to applicants in California.

📋 Description

• Directly engage with customers to address intricate technical challenges related to Kubernetes GPU clusters.

• Serve as a customer-facing Site Reliability Engineer (SRE) to ensure the health and stability of customer Kubernetes clusters.

• Become a subject matter expert on the product and act as the final technical resource before escalating issues to Engineering and Product teams.

• Monitor the health of GPU clusters and report hardware problems along with remediation strategies.

• Operate and manage production infrastructure, which includes fleet rebalancing, Slurm maintenance, node repair/migration, and Kubernetes workload oversight.

• Investigate and resolve storage and networking challenges in both bare-metal and virtual machine environments.

• Collaborate with Engineering, Research, Product, Sales, Support, and senior leadership to address customer issues and enhance customer satisfaction.

• Identify trends in support cases and translate customer feedback into actionable roadmap enhancements.

• Maintain comprehensive documentation, troubleshooting guides, procedures, and FAQs.

• Provide coverage during holidays, nights, and weekends as necessary.


⛳️ Requirements

• A minimum of 3 years of experience in a customer-facing technical position, with at least 1 year in support of an AI service or mission-critical SaaS API.

• Experience as an SRE or DevOps engineer with a focus on Kubernetes.

• Familiarity with AI, ML, GPU technologies, and high-performance computing (HPC) environments.

• Advanced understanding of Kubernetes, SLURM, Ansible, high-performance networking fabrics, NFS storage, container infrastructure, and programming/scripting languages.

• Experience in HPC/Slurm cluster environments, encompassing node draining, job scheduling, and maintenance processes.

• Knowledge of InfiniBand, RDMA, and network interface diagnostics.

• Experience with distributed storage systems like Weka and NFS, including troubleshooting I/O and bandwidth issues.

• Proficiency in installing, configuring, administering, troubleshooting, and securing compute clusters.

• Strong skills in complex technical problem-solving and proactive issue resolution.

• Capability to work collaboratively across Sales, Engineering, Support, Product, and Research teams.

• Demonstrated ownership and eagerness to acquire new skills.

• Excellent communication and interpersonal abilities, including the capacity to convey technical concepts to non-technical stakeholders.

• Ability to manage multiple projects while effectively switching contexts and prioritizing tasks.

• Willingness to work during US daytime hours, weekends, holidays, nights, and a four-day, 10-hour shift with weekend on-call duties.


🏝️ Benefits

• Equity in a startup.

• Comprehensive health insurance.

• Additional benefits.

• Flexible remote work options.

People also viewed

Immersive Gamebox4 days ago

Technical Support, Guest Experience Specialist

GB flagUnited Kingdom OnlyFull-timeSupport Engineer£30k/year
ApplyView job
Medlogix4 days ago

Application Support Analyst – MST/PST Time Zone

US flagUnited States OnlyFull-timeSupport Engineer$0 – $85k/year
ApplyView job
alt.bank4 days ago

Estagiário em Suporte Técnico N3

BR flagBrazil OnlyInternshipSupport Engineer
ApplyView job
OpenAI4 days ago

Senior Support Engineer

CA flagCanada OnlyFull-timeSupport Engineer
ApplyView job
OpenAI4 days ago

AI Support Engineer

CA flagCanada OnlyFull-timeSupport Engineer
ApplyView job
Aurum Software4 days ago

Analista de Suporte Júnior

BR flagBrazil OnlyFull-timeSupport Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers