
Senior GPU and HPC Infrastructure Engineer – DGX Cloud
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California.
• Play a vital role in a platform that streamlines GPU asset provisioning, configuration, and lifecycle management across various cloud providers.
• Develop comprehensive automation for datacenter operations, including break/fix processes and lifecycle management for extensive Machine Learning systems.
• Establish monitoring and health management features that enhance the reliability, availability, and scalability of GPU assets.
• Oversee NVLINK topology across GPU clusters.
• Create automated testing infrastructure to validate the operational capability of distributed systems.
• Work in collaboration with engineering teams to ensure seamless software integration from hardware to AI training applications.
• Over 10 years of software engineering experience working with large-scale production systems.
• Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent professional experience.
• Proficient knowledge of a systems programming language, such as Go or Python.
• Advanced understanding of Linux system administration and management.
• Familiarity with cluster management systems like Kubernetes and SLURM.
• Knowledge of performance, security, and reliability within complex distributed systems.
• Experience with system-level architecture, data synchronization, fault tolerance, and state management.
• Eligible for equity and benefits.
iRhythm Technologies, Inc.
SupplyHouse.com
Blue Ocean Global Technology
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.