Senior Site Reliability Engineer, Production Engineering

Posted 1 day ago

This is a fully remote position, open to applicants in India.

📋 Description

• Oversee a global, innovative, and cutting-edge Service Reliability Operations center.

• Provide assistance for NVIDIA Cloud products and services.

• Collaborate with Site Reliability Engineering, Security Operations Center, DevOps teams, and other departments.

• Support Production Kubernetes Services with an emphasis on automation and minimizing manual processes.

• Conduct large-scale Kubernetes administration, systems administration, and security monitoring to uphold service SLAs, integrity, and reliability.

• Utilize alerts, alarms, and observability tools to monitor, detect, prevent, and address incidents.

• Analyze logs, metrics, and system performance to troubleshoot issues.

• Lead root cause analysis and implement effective solutions.

• Initiate and facilitate incident management calls.

• Coordinate with subject matter experts and service owners for prompt incident escalation and resolution.

• Develop monitors, alarms, and alerts to enhance service reliability and customer satisfaction.


⛳️ Requirements

• More than 7 years of proven experience in administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center settings.

• Strong preference for on-premises expertise.

• Bachelor’s degree in Computer Science, Engineering, Mathematics, or equivalent experience.

• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.

• Familiarity with GPU/DPU hardware and high-performance computing cluster environments.

• Strong experience in Linux system administration, DNS, DHCP, and core Linux networking (IP Tables, routing, firewalls).

• Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure.

• Experience with CI/CD tools such as Jenkins and ArgoCD.

• Experience in scripting.

• Proficiency in programming with Python, Golang, or Rust is preferred, but not mandatory.

• Excellent communication and interpersonal skills, capable of presenting to cross-functional team members in a persuasive manner.


🏝️ Benefits

• 24/7 Production engineering team support.

• Flexibility to work on split-weekend shifts.

People also viewed

NVIDIA5 days ago

Senior Software Engineer, DGX Cloud Production Engineering

US flagCalifornia OnlyFull-timeProduction Engineer$184k – $356.5k/year
ApplyView job
RedoxSep 11

Associate Production Support Engineer, Tier I

US flagUnited States OnlyFull-timeProduction Engineer$70k – $80k/year
ApplyView job
NVIDIASep 9

Senior Storage Production Engineer – DGX Cloud

AU flagAustralia OnlyFull-timeProduction Engineer
ApplyView job
TMSSep 9

Production Support Engineer, Operations – Healthcare, Medicaid

US flagNew Jersey OnlyFreelanceProduction Engineer
ApplyView job
CanvaSep 8

Staff Production Engineer

AU flagAustralia OnlyFull-timeProduction Engineer
ApplyView job
CanvaSep 8

Senior Production Engineering Manager

AU flagAustralia OnlyFull-timeProduction Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers