
Senior Production Engineer – DGX Cloud
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Design and maintain production software, automation, and tools for control plane services, model deployments, and agentic workloads within DGX Cloud environments.
• Enhance the reliability of inference and agentic platforms and services, such as NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, alongside inference services developed with NVIDIA Dynamo.
• Boost availability of endpoints, manage inference routing, oversee capacity, and ensure service health.
• Implement infrastructure as code and GitOps methodologies to deploy, configure, validate, upgrade, and restore services uniformly across various environments.
• Create workflows for service activation, model releases, handoffs, deprecation, and ongoing operational tasks.
• Establish and monitor SLIs and SLOs for inference and control plane services, using error budgets to inform reliability enhancements.
• Engage in on-call duties and incident response, troubleshoot failures, and transform recurrent issues into automated solutions and long-lasting fixes.
• Work in collaboration with model, platform, storage, networking, security, and GPU infrastructure teams to design and manage services safely at scale.
• Over 8 years of experience in building or managing production services and large-scale distributed systems, including practical automation experience.
• Proficient programming skills in Python, Go, or a similar language.
• Experience in developing tools for production operations.
• Familiarity with infrastructure as code, configuration management, or GitOps practices.
• Proven experience in creating automation for consistent service deployments and modifications.
• Strong understanding of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking basics.
• Capability to diagnose production failures effectively.
• Knowledge of SRE principles, including SLIs, SLOs, error budgets, incident response, and minimizing operational toil.
• Experience in instrumenting services and utilizing metrics, logs, and traces to analyze system behavior and enhance reliability.
• Excellent technical communication skills and the ability to collaborate across engineering teams.
• BS/MS in Computer Science or equivalent experience.
• Equity
• Benefits
NVIDIA
NVIDIA
SAIC
SAIC
Get handpicked remote jobs straight to your inbox weekly.