
Technical Support Engineer – GPU Clusters
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in California.
• Directly engage with customers to address intricate technical challenges related to Kubernetes GPU clusters.
• Serve as a customer-facing Site Reliability Engineer (SRE) to ensure the health and stability of customer Kubernetes clusters.
• Become a subject matter expert on the product and act as the final technical resource before escalating issues to Engineering and Product teams.
• Monitor the health of GPU clusters and report hardware problems along with remediation strategies.
• Operate and manage production infrastructure, which includes fleet rebalancing, Slurm maintenance, node repair/migration, and Kubernetes workload oversight.
• Investigate and resolve storage and networking challenges in both bare-metal and virtual machine environments.
• Collaborate with Engineering, Research, Product, Sales, Support, and senior leadership to address customer issues and enhance customer satisfaction.
• Identify trends in support cases and translate customer feedback into actionable roadmap enhancements.
• Maintain comprehensive documentation, troubleshooting guides, procedures, and FAQs.
• Provide coverage during holidays, nights, and weekends as necessary.
• A minimum of 3 years of experience in a customer-facing technical position, with at least 1 year in support of an AI service or mission-critical SaaS API.
• Experience as an SRE or DevOps engineer with a focus on Kubernetes.
• Familiarity with AI, ML, GPU technologies, and high-performance computing (HPC) environments.
• Advanced understanding of Kubernetes, SLURM, Ansible, high-performance networking fabrics, NFS storage, container infrastructure, and programming/scripting languages.
• Experience in HPC/Slurm cluster environments, encompassing node draining, job scheduling, and maintenance processes.
• Knowledge of InfiniBand, RDMA, and network interface diagnostics.
• Experience with distributed storage systems like Weka and NFS, including troubleshooting I/O and bandwidth issues.
• Proficiency in installing, configuring, administering, troubleshooting, and securing compute clusters.
• Strong skills in complex technical problem-solving and proactive issue resolution.
• Capability to work collaboratively across Sales, Engineering, Support, Product, and Research teams.
• Demonstrated ownership and eagerness to acquire new skills.
• Excellent communication and interpersonal abilities, including the capacity to convey technical concepts to non-technical stakeholders.
• Ability to manage multiple projects while effectively switching contexts and prioritizing tasks.
• Willingness to work during US daytime hours, weekends, holidays, nights, and a four-day, 10-hour shift with weekend on-call duties.
• Equity in a startup.
• Comprehensive health insurance.
• Additional benefits.
• Flexible remote work options.
Immersive Gamebox
Medlogix
alt.bank
Get handpicked remote jobs straight to your inbox weekly.