
Senior Site Reliability Engineer, BCM – DGX Cloud
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Participate in the implementation and routine management of extensive next-generation GPU platforms.
• Address incidents within GPU clusters, facilitating communication between cluster operations and development teams.
• Create and implement minor features within the Base Command Manager product.
• Assess intricate cluster configurations, including Slurm and Kubernetes orchestrators, focusing on performance, scalability, and resilience.
• Guarantee that cluster configurations align with authentic customer scenarios.
• Provide support to external customers utilizing NVIDIA solution clusters, as well as to internal research, operations, and next-generation project clusters.
• Bachelor's Degree or equivalent experience in Computer Science or a related discipline.
• Over 8 years of experience in site reliability engineering and/or software development positions.
• Proficient in Python.
• Strong understanding of Linux and networking principles.
• Experience with C++, high-performance computing, Kubernetes, and/or system administration is advantageous.
• Prior experience managing BCM/Bright Cluster Manager/Base Command Manager clusters is beneficial.
• Expertise in cluster networking, including InfiniBand and Spectrum-X.
• Equity.
• Comprehensive benefits.
• An inclusive and supportive work environment.
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.