
Senior Software Engineer, DGX Cloud Production Engineering
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in California.
β’ Design and manage automation for extensive GPU clusters within NVIDIA Cloud Partners and on-premises environments.
β’ Create tools and services for provisioning, validation, upgrades, monitoring, repair, and overall cluster lifecycle management.
β’ Enhance Day 0, Day 1, and Day 2 workflows for cluster setup, transition, and operational excellence.
β’ Minimize manual interventions in production through APIs, GitOps, automation, and agent-assisted processes.
β’ Engage in on-call duties, incident response, debugging, and thorough follow-up activities.
β’ Collaborate with platform, storage, networking, security, and workload teams to ensure infrastructure is production-ready.
β’ Over 8 years of experience in building or operating production infrastructure.
β’ Proficient programming skills in Python, Go, or similar languages.
β’ Familiarity with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
β’ Capability to troubleshoot distributed systems in production environments.
β’ Strong communication skills and the ability to collaborate across teams.
β’ BS/MS in Computer Science or equivalent experience.
β’ Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.
β’ Knowledge of SLOs, on-call responsibilities, incident response, observability, and reliability practices.
β’ Exposure to BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure.
β’ Equity
β’ Benefits
SAIC
SAIC
NVIDIA
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.