
Senior Software Engineer, DGX Cloud Production Engineering
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California.
• Design and manage automation for extensive GPU clusters within NVIDIA Cloud Partners and on-premises settings.
• Create tools and services for provisioning, validation, upgrades, monitoring, repairs, and managing cluster lifecycle operations.
• Enhance workflows for Day 0, Day 1, and Day 2 processes related to cluster setup, handoff, and operational production.
• Minimize manual interventions in production through APIs, GitOps, automation, and agent-assisted processes.
• Engage in on-call duties, incident response, debugging, and thorough follow-up tasks.
• Collaborate with platform, storage, networking, security, and workload teams to ensure infrastructure is ready for production.
• A minimum of 8 years of experience in building or managing production infrastructure.
• Proficient programming skills in Python, Go, or a related language.
• Familiarity with Linux, Kubernetes, containers, cloud infrastructure, or automation of infrastructure.
• Capability to troubleshoot distributed systems in a production environment.
• Excellent communication skills and ability to collaborate effectively across teams.
• A BS/MS degree in Computer Science or equivalent professional experience.
• Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.
• Knowledge of SLOs, on-call responsibilities, incident response, observability, and reliability practices.
• Background in BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure.
• Equity
• Benefits
Redox
NVIDIA
TMS
Get handpicked remote jobs straight to your inbox weekly.