
Principal Software Engineer, DGX Cloud Production Engineering
Posted Aug 12

Posted Aug 12
This is a fully remote position, open to applicants in California.
β’ Oversee the design and development of essential Kubernetes platform functionalities, encompassing cluster management, control plane services, fleet lifecycle, and day-2 operations.
β’ Create and develop highly dependable distributed systems and APIs for the provisioning, management, upgrading, and remediation of Kubernetes clusters at scale.
β’ Establish technical requirements, validation standards, production-readiness practices, and guidance for declarative workflows and automation throughout the Kubernetes ecosystem.
β’ Work collaboratively with engineering teams to craft cohesive platform experiences that integrate management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
β’ Spearhead the identification and resolution of intricate platform challenges involving infrastructure, runtime, networking, hardware, and operations.
β’ Enhance the scalability, resilience, and operability of systems that underpin large-scale AI deployments.
β’ Shape engineering standards, architectural choices, and the long-term strategy for the platform.
β’ Guide senior engineers and elevate the standards for design quality, execution, and engineering rigor throughout the organization.
β’ Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related discipline, or equivalent professional experience.
β’ 15+ years of pertinent software engineering experience, particularly in the development and operation of large-scale production systems.
β’ In-depth knowledge of Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
β’ Strong foundation in distributed systems design, reliability, scalability, and failure recovery.
β’ Documented experience in developing platform software, infrastructure control planes, or foundations for managed services.
β’ Proficient programming skills in one or more systems or cloud-native languages, including Go, Python, Rust, or C++.
β’ Experience in designing clear APIs and abstractions for platform users and engineering teams.
β’ Capability to provide technical leadership across various teams and to drive complex, cross-functional initiatives to successful completion.
β’ Exceptional communication and collaboration abilities, supported by significant technical contributions and recognized expertise in influencing departmental architecture and high-priority company initiatives.
β’ Experience developing Kubernetes platforms or managed Kubernetes services.
β’ Expertise in fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations.
β’ Familiarity with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation.
β’ Understanding of public-cloud and bare-metal infrastructure environments.
β’ Experience in supporting AI, GPU, HPC, or other large-scale accelerated computing platforms.
β’ Competitive salaries
β’ Generous benefits package
β’ Equity
β’ Additional benefits
NVIDIA
SAIC
SAIC
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.