
Senior Software Engineer, Resilience Engineering - DGX Cloud
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in California.
• Develop a comprehensive reliability strategy for the organization, steering NVIDIA's operational practices in a 24/7 setting.
• Establish a robust Service Level Objective (SLO) program, ensuring high standards are defined and upheld across various teams.
• Oversee incident management for high-severity situations, focusing on achieving efficient resolutions with minimal disruption.
• Enhance and refine production code on a daily basis, improving our data platform and associated tools.
• Execute chaos engineering, failure injection, and resilience testing to elevate our team's standard operating procedures.
• Set a benchmark for quality by demonstrating your hands-on experience and leadership qualities.
• A minimum of 8 years of experience in the industry.
• Bachelor's or Master's degree, or equivalent experience in managing operating systems at scale.
• Strong software engineering background with practical, hands-on experience in Go, Python, or similar languages.
• Demonstrated track record in establishing and maintaining a rigorous SLO program.
• Practical knowledge in reliability disciplines such as chaos engineering and failure injection.
• Ability to influence and collaborate effectively across team boundaries, leveraging credibility and expertise.
• Equity
• Comprehensive benefits package
Sigma Software Group
Collectly
Allata
Get handpicked remote jobs straight to your inbox weekly.