Senior Production Engineer – DGX Cloud

atNVIDIARemoteUS flagCaliforniaFull-timeProduction EngineerSenior$184k – $356.5k/year

Posted 2 days ago

This is a fully remote position, open to applicants in California.

📋 Description

• Design and maintain production software, automation, and tools for control plane services, model deployments, and agentic workloads within DGX Cloud environments.

• Enhance the reliability of inference and agentic platforms and services, such as NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, alongside inference services developed with NVIDIA Dynamo.

• Boost availability of endpoints, manage inference routing, oversee capacity, and ensure service health.

• Implement infrastructure as code and GitOps methodologies to deploy, configure, validate, upgrade, and restore services uniformly across various environments.

• Create workflows for service activation, model releases, handoffs, deprecation, and ongoing operational tasks.

• Establish and monitor SLIs and SLOs for inference and control plane services, using error budgets to inform reliability enhancements.

• Engage in on-call duties and incident response, troubleshoot failures, and transform recurrent issues into automated solutions and long-lasting fixes.

• Work in collaboration with model, platform, storage, networking, security, and GPU infrastructure teams to design and manage services safely at scale.


⛳️ Requirements

• Over 8 years of experience in building or managing production services and large-scale distributed systems, including practical automation experience.

• Proficient programming skills in Python, Go, or a similar language.

• Experience in developing tools for production operations.

• Familiarity with infrastructure as code, configuration management, or GitOps practices.

• Proven experience in creating automation for consistent service deployments and modifications.

• Strong understanding of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking basics.

• Capability to diagnose production failures effectively.

• Knowledge of SRE principles, including SLIs, SLOs, error budgets, incident response, and minimizing operational toil.

• Experience in instrumenting services and utilizing metrics, logs, and traces to analyze system behavior and enhance reliability.

• Excellent technical communication skills and the ability to collaborate across engineering teams.

• BS/MS in Computer Science or equivalent experience.


🏝️ Benefits

• Equity

• Benefits

People also viewed

NVIDIA5 days ago

Senior Software Engineer, SRE, Production Engineering

US flagCalifornia OnlyFull-timeProduction Engineer$152k – $287.5k/year
ApplyView job
NVIDIASep 23

Senior Software Engineer, DGX Cloud Production Engineering

US flagCalifornia OnlyFull-timeProduction Engineer$184k – $356.5k/year
ApplyView job
SAICSep 21

Senior Systems Engineer – Production Support

US flagTexas OnlyFull-timeProduction Engineer
ApplyView job
SAICSep 21

Senior Systems Engineer – Production Support

US flagTexas OnlyFull-timeProduction Engineer
ApplyView job
NVIDIASep 21

Senior Site Reliability Engineer, Production Engineering

IN flagIndia OnlyFull-timeProduction Engineer
ApplyView job
NVIDIASep 16

Senior Software Engineer, DGX Cloud Production Engineering

US flagCalifornia OnlyFull-timeProduction Engineer$184k – $356.5k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers