Senior Systems Software Engineer, Kubernetes Scale – DGX Cloud

Posted 2 days ago

This is a fully remote position, open to applicants in Poland, +5 more countries.

📋 Description

• Lead the comprehensive performance and scale evaluation for the NVIDIA DGX Cloud software stack, covering Kubernetes control and data planes, GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving.

• Work closely with AI researchers, developers, and clients to create automated tests that emulate real user workloads.

• Analyze performance and scalability challenges in intricate distributed systems and determine their root causes.

• Design and construct monitoring, reporting, and analysis tools for performance and scale testing across software, GPU, and CPU resources.

• Diagnose, troubleshoot, and resolve issues associated with managing Kubernetes clusters at an exceptionally large scale.

• Develop and sustain a continuous framework for performance and scale testing utilizing a modern CI/CD pipeline.

• Record research, methodologies, and outcomes, and share insights at both internal and external forums, including KubeCon and GTC.

• Participate in Kubernetes, CNCF, and NVIDIA open-source communities to validate AI workload performance and influence design choices.


⛳️ Requirements

• Over 8 years of experience in Computer Architecture, Networking, Storage systems, and Accelerators.

• Bachelor's or Master's degree in Engineering, ideally in Electrical Engineering, Computer Engineering, or Computer Science, or equivalent experience.

• Proficient in Kubernetes with knowledge of related CNCF projects.

• Experience with large-scale parallel and distributed accelerator-based systems.

• Expertise in optimizing performance and AI workloads on extensive systems.

• Familiarity with performance modeling and benchmarking at scale.

• Proficient in Golang and Python.

• Familiarity with the NVIDIA software ecosystem in training and inference domains.

• Expertise with at least one public CSP infrastructure, such as GCP, AWS, Azure, or OCI.

• Strong operational experience with a Kubernetes distribution.

• Previous experience in scaling Kubernetes clusters to ultra-large node and object counts.

• Proven track record of contributing to the open-source community.

• PhD in relevant fields (highly regarded qualification).

• Exceptional communication and interpersonal skills.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work opportunities.

• Generous paid time off and holidays.

• Opportunities for professional development and continuing education.

People also viewed

Smartitory1 day ago

Java Fejlesztő

RO flagRomania OnlyFull-timeBackend Engineer
ApplyView job
Händlerbund1 day ago

Backend Developer, PHP, Laravel

DE flagGermany OnlyFull-timeBackend Engineer
ApplyView job
ASRC Federal1 day ago

Senior Drupal Subject Matter Expert

US flagDistrict of Columbia, +1 more stateFull-timeBackend Engineer
ApplyView job
SPD Technology1 day ago

Senior Backend Engineer – Payments, Go, TypeScript, Node.js

UA flagUkraine, +5 more countriesFull-timeBackend Engineer
ApplyView job
BigDataCorp1 day ago

Mid-Level Back-End Developer – Research and Development

BR flagBrazil OnlyFull-timeBackend Engineer
ApplyView job
Modash1 day ago

Senior Backend Engineer, Data Core

EE flagEstonia, +5 more countriesFull-timeBackend Engineer€80k – €110k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers