
Principal Software Engineer, DGX Cloud Production Engineering
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Oversee the design and development of essential Kubernetes platform functionalities, encompassing cluster management, control plane services, fleet lifecycle, and day-2 operations.
• Create and develop highly dependable distributed systems and APIs for the provisioning, management, upgrading, and remediation of Kubernetes clusters at scale.
• Establish technical requirements, validation standards, production-readiness practices, and guidance for declarative workflows and automation throughout the Kubernetes ecosystem.
• Work collaboratively with engineering teams to craft cohesive platform experiences that integrate management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
• Spearhead the identification and resolution of intricate platform challenges involving infrastructure, runtime, networking, hardware, and operations.
• Enhance the scalability, resilience, and operability of systems that underpin large-scale AI deployments.
• Shape engineering standards, architectural choices, and the long-term strategy for the platform.
• Guide senior engineers and elevate the standards for design quality, execution, and engineering rigor throughout the organization.
• Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related discipline, or equivalent professional experience.
• 15+ years of pertinent software engineering experience, particularly in the development and operation of large-scale production systems.
• In-depth knowledge of Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
• Strong foundation in distributed systems design, reliability, scalability, and failure recovery.
• Documented experience in developing platform software, infrastructure control planes, or foundations for managed services.
• Proficient programming skills in one or more systems or cloud-native languages, including Go, Python, Rust, or C++.
• Experience in designing clear APIs and abstractions for platform users and engineering teams.
• Capability to provide technical leadership across various teams and to drive complex, cross-functional initiatives to successful completion.
• Exceptional communication and collaboration abilities, supported by significant technical contributions and recognized expertise in influencing departmental architecture and high-priority company initiatives.
• Experience developing Kubernetes platforms or managed Kubernetes services.
• Expertise in fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations.
• Familiarity with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation.
• Understanding of public-cloud and bare-metal infrastructure environments.
• Experience in supporting AI, GPU, HPC, or other large-scale accelerated computing platforms.
• Competitive salaries
• Generous benefits package
• Equity
• Additional benefits
Vultr
Canva
Canva
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.