
Senior Software Engineer, DGX Cloud Production Engineering
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Develop the next generation of NVIDIA's Kubernetes platform.
• Collaborate on essential Kubernetes platform functionalities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
• Design and create dependable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
• Establish technical requirements, validation criteria, production-readiness practices, and guidance for declarative workflows and automation.
• Craft cohesive platform experiences that encompass management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
• Diagnose and resolve intricate platform issues across infrastructure, runtime, networking, hardware, and operations.
• Enhance scalability, resilience, and operability for systems supporting large-scale AI deployments.
• Offer technical leadership across team boundaries and drive cross-functional initiatives to successful completion.
• BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
• Over 5 years of relevant software engineering experience, including the development and operation of large-scale production systems.
• In-depth knowledge of Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
• Strong foundation in distributed systems design, reliability, scalability, and failure recovery.
• Experience in building platform software, infrastructure control planes, or managed-service foundations.
• Proficient programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++.
• Experience in designing clear APIs and abstractions for platform consumers and engineering teams.
• Capability to provide technical leadership across team boundaries and navigate ambiguous, cross-functional initiatives to completion.
• Exceptional communication and collaboration skills, with significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives.
• Background in Kubernetes platforms or managed Kubernetes services.
• Experience in fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations.
• Familiarity with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation.
• Awareness of public-cloud and bare-metal infrastructure environments.
• Experience supporting AI, GPU, HPC, or other large-scale accelerated computing platforms.
• Competitive salaries.
• Generous benefits package.
• Equity.
• Additional benefits.
Natera
Palo Alto Networks
EXL
Vultr
Get handpicked remote jobs straight to your inbox weekly.