Principal Software Engineer, Distributed Systems Engineer – DGX Cloud

atNVIDIARemoteUS flagNorth CarolinaFull-timeBlockchain EngineerLead$272k – $431.3k/year

Posted 5 days ago

This is a fully remote position, open to applicants in North Carolina.

📋 Description

• Play a vital role in the DGX Cloud team, which manages production systems that support large-scale GPU clusters for AI workloads.

• Create tailored software solutions for scheduling GPU resources on Kubernetes.

• Establish monitoring and health-management features to ensure the reliability, availability, and scalability of GPU assets.

• Leverage data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry.

• Collaborate with various teams at NVIDIA to maintain the reliable, consistent, and high-performance operation of production AI clusters.

• Assess system failures and enhance services through a structured incident-management approach.


⛳️ Requirements

• Extensive software engineering experience with Kubernetes, encompassing cluster operations, operator development, node health monitoring, and GPU resource scheduling.

• Proven experience in a software engineering position within a highly technical environment, demonstrating measurable impact.

• Software development expertise with Kubernetes APIs and frameworks, beyond just cluster operation.

• Excellent communication abilities and capacity to collaborate with multifunctional teams, principals, and architects across different organizational levels and locations.

• Over 15 years in a similar position with experience in large-scale production systems.

• Familiarity with standard software engineering principles, tools, and methodologies.

• Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.

• Proficiency in systems programming languages, including Go or Python.

• Strong understanding of data structures and algorithms.

• Technical expertise in managing and automating large-scale distributed systems, regardless of cloud provider.

• Advanced practical experience and thorough knowledge of cluster management systems like Kubernetes, Slurm, or Bright Cluster Manager.

• Demonstrated operational excellence in maintaining reliable and efficient AI infrastructure.


🏝️ Benefits

• Equity

• Benefits

People also viewed

NVIDIASep 28

Senior Software Engineer, Distributed Systems Engineer

US flagCalifornia OnlyFull-timeBlockchain Engineer$152k – $287.5k/year
ApplyView job
HighLevelSep 25

Staff Engineer – Distributed Systems

IN flagIndia OnlyFull-timeBlockchain Engineer
ApplyView job
CoinbaseSep 24

Senior Software Engineer – Blockchain Platform, Wallets, Liquidity, Bridging

US flagUnited States OnlyFull-timeBlockchain Engineer$186.1k – $218.9k/year
ApplyView job
Tether.toSep 24

Blockchain Investigator, Law Enforcement Liaison

NZ flagNew Zealand OnlyFull-timeBlockchain Engineer
ApplyView job
NetflixSep 22

Distributed Systems Engineer 5 – Pricing, Catalog, and Offers

US flagColorado OnlyFull-timeBlockchain Engineer$349.2k – $557.1k/year
ApplyView job
MoznSep 21

Distributed Systems Engineer III

EG flagEgypt OnlyFull-timeBlockchain Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers