
Principal Software Engineer, Distributed Systems Engineer – DGX Cloud
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in North Carolina.
• Play a vital role in the DGX Cloud team, which manages production systems that support large-scale GPU clusters for AI workloads.
• Create tailored software solutions for scheduling GPU resources on Kubernetes.
• Establish monitoring and health-management features to ensure the reliability, availability, and scalability of GPU assets.
• Leverage data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry.
• Collaborate with various teams at NVIDIA to maintain the reliable, consistent, and high-performance operation of production AI clusters.
• Assess system failures and enhance services through a structured incident-management approach.
• Extensive software engineering experience with Kubernetes, encompassing cluster operations, operator development, node health monitoring, and GPU resource scheduling.
• Proven experience in a software engineering position within a highly technical environment, demonstrating measurable impact.
• Software development expertise with Kubernetes APIs and frameworks, beyond just cluster operation.
• Excellent communication abilities and capacity to collaborate with multifunctional teams, principals, and architects across different organizational levels and locations.
• Over 15 years in a similar position with experience in large-scale production systems.
• Familiarity with standard software engineering principles, tools, and methodologies.
• Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
• Proficiency in systems programming languages, including Go or Python.
• Strong understanding of data structures and algorithms.
• Technical expertise in managing and automating large-scale distributed systems, regardless of cloud provider.
• Advanced practical experience and thorough knowledge of cluster management systems like Kubernetes, Slurm, or Bright Cluster Manager.
• Demonstrated operational excellence in maintaining reliable and efficient AI infrastructure.
• Equity
• Benefits
NVIDIA
HighLevel
Coinbase
Tether.to
Get handpicked remote jobs straight to your inbox weekly.