
Senior Software Engineer, Distributed Systems Engineer
Posted Sep 28

Posted Sep 28
This is a fully remote position, open to applicants in California.
• Join the DGX Cloud team tasked with managing production systems that support extensive GPU clusters for AI applications.
• Create tailored software solutions for scheduling GPU resources on Kubernetes.
• Establish monitoring and health management features to ensure the reliability, availability, and scalability of GPU resources.
• Utilize data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry.
• Work collaboratively with various teams within NVIDIA to guarantee that production AI clusters operate reliably, consistently, and at peak performance.
• Analyze system failures and enhance services through a structured incident management process.
• Extensive software engineering experience with Kubernetes, including cluster operations, operator development, node health monitoring, and GPU resource scheduling.
• Proven experience in a software engineering position within a highly technical environment, showcasing measurable impact.
• Experience in software development with Kubernetes APIs and frameworks, rather than merely managing a cluster.
• Excellent communication skills with the ability to collaborate effectively with multifunctional teams, principals, and architects across different organizational boundaries and locations.
• Over 5 years of experience in a similar role with a focus on large-scale production systems.
• Familiarity with standard software engineering principles, tools, and methodologies.
• Bachelor’s degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent work experience.
• Proficiency in systems programming languages such as Go or Python.
• Strong grasp of data structures and algorithms.
• Technical expertise in managing and automating large-scale distributed systems, independent of cloud providers.
• Advanced practical experience and a deep understanding of Kubernetes, Slurm, or Bright Cluster Manager.
• Demonstrated operational excellence in maintaining reliable and high-performing AI infrastructure.
• Equity
• Benefits
NVIDIA
HighLevel
Coinbase
Tether.to
Get handpicked remote jobs straight to your inbox weekly.