
Senior Software Engineer – Distributed Systems
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in California, +2 more states.
• Creating and building a highly distributed and scalable platform designed to identify, diagnose, and remediate underperforming GPU assets.
• Collaborating with various teams within NVIDIA to guarantee that production AI clusters operate reliably, consistently, and at peak performance.
• Assessing system failures and enhancing services according to a clearly defined incident management process.
• Joining a DGX Cloud team that oversees production systems facilitating large-scale GPU clusters for diverse AI workloads.
• Over 5 years of experience in a similar role with a background in large-scale production systems.
• Proven experience in a software engineering position within a highly technical organization, showcasing the impact of your contributions.
• Technical expertise, including proficiency in a systems programming language (Go, Python), along with a solid grasp of data structures and algorithms.
• Bachelor’s degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
• Highly driven with excellent communication skills, capable of successfully collaborating with cross-functional teams, principles, and architects while effectively coordinating across organizational boundaries and locations.
• Equity
• Benefits
Coinbase
DMS International
Netflix
Netflix
Get handpicked remote jobs straight to your inbox weekly.