
Senior Site Reliability Engineer, Production Engineering
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Oversee a global, innovative, and cutting-edge Service Reliability Operations center.
• Provide assistance for NVIDIA Cloud products and services.
• Collaborate with Site Reliability Engineering, Security Operations Center, DevOps teams, and other departments.
• Support Production Kubernetes Services with an emphasis on automation and minimizing manual processes.
• Conduct large-scale Kubernetes administration, systems administration, and security monitoring to uphold service SLAs, integrity, and reliability.
• Utilize alerts, alarms, and observability tools to monitor, detect, prevent, and address incidents.
• Analyze logs, metrics, and system performance to troubleshoot issues.
• Lead root cause analysis and implement effective solutions.
• Initiate and facilitate incident management calls.
• Coordinate with subject matter experts and service owners for prompt incident escalation and resolution.
• Develop monitors, alarms, and alerts to enhance service reliability and customer satisfaction.
• More than 7 years of proven experience in administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center settings.
• Strong preference for on-premises expertise.
• Bachelor’s degree in Computer Science, Engineering, Mathematics, or equivalent experience.
• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
• Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
• Strong experience in Linux system administration, DNS, DHCP, and core Linux networking (IP Tables, routing, firewalls).
• Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure.
• Experience with CI/CD tools such as Jenkins and ArgoCD.
• Experience in scripting.
• Proficiency in programming with Python, Golang, or Rust is preferred, but not mandatory.
• Excellent communication and interpersonal skills, capable of presenting to cross-functional team members in a persuasive manner.
• 24/7 Production engineering team support.
• Flexibility to work on split-weekend shifts.
NVIDIA
Redox
NVIDIA
TMS
Get handpicked remote jobs straight to your inbox weekly.