
Senior Engineer, NCX
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in Germany, +2 more countries.
• Spearhead the operational readiness initiatives for NVIDIA Cloud Partner Day 2.
• Work in conjunction with NVIDIA Cloud Partners to create systems, procedures, automation, and operational methodologies for overseeing NVIDIA accelerated infrastructure following deployment and activation.
• Formulate continuous validation strategies for GPU, CPU, storage, and network health across expansive AI clusters.
• Set up telemetry, monitoring, alerting, dashboards, and operational signals throughout compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.
• Design automated workflows to identify, isolate, drain, repair, validate, and reinstate unhealthy infrastructure into service.
• Execute scalable strategies for GPU fleet lifecycle management, encompassing driver and firmware management, Kubernetes node upkeep, OS patching, configuration management, upgrades, and detection of configuration drift.
• Convert NVIDIA NCP specifications and reference architectures into operational practices, validation benchmarks, runbooks, automation, and quantifiable operational standards.
• Create health indicators, SLOs, metrics, acceptance criteria, and ongoing validation processes for infrastructure reliability and service readiness.
• Produce reusable tools, automation scripts, implementation guides, runbooks, operational playbooks, and reference implementations across various NCP environments.
• Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent experience.
• Over 8 years of expertise in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or comparable roles within large-scale production settings.
• Significant experience managing Linux-based distributed systems and cloud infrastructure in a production environment.
• Profound knowledge of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node configurations.
• Comprehensive understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations governed by service level agreements.
• Experience in developing automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
• Solid networking principles and experience diagnosing complex distributed systems across compute, network, and storage layers.
• Programming and automation proficiency using Python, Go, shell scripting, or similar languages.
• Experience in managing large-scale GPU or accelerated computing infrastructure that supports AI training and inference tasks.
• Familiarity with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.
• Demonstrated experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure while operating SLOs for large-scale compute infrastructure and leveraging operational data to enhance availability, performance, and fleet efficiency.
• In-depth knowledge of infrastructure observability tools like Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.
• Awareness of failure modes concerning large distributed AI workloads and the essential infrastructure features required to reliably support extended training and production inference.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work opportunities.
• Professional development and training programs.
• Generous vacation and paid time off policy.
TMS
StaffMill Oy
HumanIT Digital Consulting
HumanIT Digital Consulting
Get handpicked remote jobs straight to your inbox weekly.