Senior Engineer, NCX

Posted Sep 1

This is a fully remote position, open to applicants in Germany, +2 more countries.

📋 Description

• Spearhead the operational readiness initiatives for NVIDIA Cloud Partner Day 2.

• Work in conjunction with NVIDIA Cloud Partners to create systems, procedures, automation, and operational methodologies for overseeing NVIDIA accelerated infrastructure following deployment and activation.

• Formulate continuous validation strategies for GPU, CPU, storage, and network health across expansive AI clusters.

• Set up telemetry, monitoring, alerting, dashboards, and operational signals throughout compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.

• Design automated workflows to identify, isolate, drain, repair, validate, and reinstate unhealthy infrastructure into service.

• Execute scalable strategies for GPU fleet lifecycle management, encompassing driver and firmware management, Kubernetes node upkeep, OS patching, configuration management, upgrades, and detection of configuration drift.

• Convert NVIDIA NCP specifications and reference architectures into operational practices, validation benchmarks, runbooks, automation, and quantifiable operational standards.

• Create health indicators, SLOs, metrics, acceptance criteria, and ongoing validation processes for infrastructure reliability and service readiness.

• Produce reusable tools, automation scripts, implementation guides, runbooks, operational playbooks, and reference implementations across various NCP environments.


⛳️ Requirements

• Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent experience.

• Over 8 years of expertise in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or comparable roles within large-scale production settings.

• Significant experience managing Linux-based distributed systems and cloud infrastructure in a production environment.

• Profound knowledge of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node configurations.

• Comprehensive understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations governed by service level agreements.

• Experience in developing automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.

• Solid networking principles and experience diagnosing complex distributed systems across compute, network, and storage layers.

• Programming and automation proficiency using Python, Go, shell scripting, or similar languages.

• Experience in managing large-scale GPU or accelerated computing infrastructure that supports AI training and inference tasks.

• Familiarity with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.

• Demonstrated experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure while operating SLOs for large-scale compute infrastructure and leveraging operational data to enhance availability, performance, and fleet efficiency.

• In-depth knowledge of infrastructure observability tools like Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.

• Awareness of failure modes concerning large distributed AI workloads and the essential infrastructure features required to reliably support extended training and production inference.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work opportunities.

• Professional development and training programs.

• Generous vacation and paid time off policy.

People also viewed

TMS1 day ago

.Net Full Stack Lead, Kentico CMS, Genesys Cloud

US flagCalifornia OnlyFreelanceFull-stack Engineer
ApplyView job
StaffMill Oy1 day ago

Junior Software Developer

FI flagFinland OnlyPart-timeFull-stack Engineer
ApplyView job
HumanIT Digital Consulting1 day ago

Senior/Staff Java Fullstack Developer – Java, Spring Boot, React

PT flagPortugal OnlyFull-timeFull-stack Engineer
ApplyView job
HumanIT Digital Consulting1 day ago

Fullstack Developer, Java, React, AI Tooling

PT flagPortugal OnlyFull-timeFull-stack Engineer€2,550 – €3,350/month
ApplyView job
Humana1 day ago

Manager, Full Stack Engineering

US flagFlorida, +8 more statesFull-timeFull-stack Engineer$117.6k – $161.7k/year
ApplyView job
Capital Technology Group, LLC1 day ago

Software Developer

US flagUnited States OnlyFull-timeFull-stack Engineer$75k – $110k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers