
Senior Systems Software Engineer, Kubernetes Scale – DGX Cloud
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Poland, +5 more countries.
• Lead the comprehensive performance and scale evaluation for the NVIDIA DGX Cloud software stack, covering Kubernetes control and data planes, GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving.
• Work closely with AI researchers, developers, and clients to create automated tests that emulate real user workloads.
• Analyze performance and scalability challenges in intricate distributed systems and determine their root causes.
• Design and construct monitoring, reporting, and analysis tools for performance and scale testing across software, GPU, and CPU resources.
• Diagnose, troubleshoot, and resolve issues associated with managing Kubernetes clusters at an exceptionally large scale.
• Develop and sustain a continuous framework for performance and scale testing utilizing a modern CI/CD pipeline.
• Record research, methodologies, and outcomes, and share insights at both internal and external forums, including KubeCon and GTC.
• Participate in Kubernetes, CNCF, and NVIDIA open-source communities to validate AI workload performance and influence design choices.
• Over 8 years of experience in Computer Architecture, Networking, Storage systems, and Accelerators.
• Bachelor's or Master's degree in Engineering, ideally in Electrical Engineering, Computer Engineering, or Computer Science, or equivalent experience.
• Proficient in Kubernetes with knowledge of related CNCF projects.
• Experience with large-scale parallel and distributed accelerator-based systems.
• Expertise in optimizing performance and AI workloads on extensive systems.
• Familiarity with performance modeling and benchmarking at scale.
• Proficient in Golang and Python.
• Familiarity with the NVIDIA software ecosystem in training and inference domains.
• Expertise with at least one public CSP infrastructure, such as GCP, AWS, Azure, or OCI.
• Strong operational experience with a Kubernetes distribution.
• Previous experience in scaling Kubernetes clusters to ultra-large node and object counts.
• Proven track record of contributing to the open-source community.
• PhD in relevant fields (highly regarded qualification).
• Exceptional communication and interpersonal skills.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work opportunities.
• Generous paid time off and holidays.
• Opportunities for professional development and continuing education.
Händlerbund
ASRC Federal
SPD Technology
Get handpicked remote jobs straight to your inbox weekly.