
Senior Engineer
Posted Aug 24

Posted Aug 24
This is a fully remote position, open to applicants in California, +1 more state.
β’ Oversee NVIDIA Cloud Partner Day 2 operational readiness initiatives post-initial deployment and activation.
β’ Partner with NVIDIA Cloud Partners to create systems, procedures, automation, and operational methodologies for optimized infrastructure.
β’ Develop ongoing validation for GPU, CPU, storage, and network health within extensive AI clusters.
β’ Set up telemetry, monitoring, alerting, dashboards, and operational indicators across compute, GPU, networking, storage, Kubernetes, and AI workloads.
β’ Create automated workflows to detect, isolate, drain, repair, validate, and restore non-functional infrastructure back to service.
β’ Oversee the GPU fleet lifecycle, encompassing drivers, firmware, Kubernetes nodes, OS updates, configuration management, upgrades, and configuration drift.
β’ Convert NVIDIA NCP requirements and reference architectures into production operational practices, validation standards, runbooks, automation, and measurable benchmarks.
β’ Define health indicators, SLOs, metrics, acceptance standards, and infrastructure readiness validation.
β’ Develop reusable tools, automation, implementation guides, runbooks, playbooks, and reference implementations throughout NCP environments.
β’ BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent experience.
β’ Over 8 years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or comparable roles in large-scale production settings.
β’ Extensive experience in managing Linux-based distributed systems and cloud infrastructure in a production environment.
β’ Profound knowledge of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.
β’ Strong grasp of production observability, including metrics, logging, alerting, dashboards, health checks, and operations driven by service level agreements.
β’ Experience in creating automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
β’ Solid networking principles and experience troubleshooting complex distributed systems across compute, network, and storage layers.
β’ Programming and automation skills using Python, Go, shell scripting, or similar languages.
β’ Experience managing large GPU or accelerated computing infrastructure that supports AI training and inference workloads.
β’ Familiarity with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.
β’ Proven track record of collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure while operating SLOs for large-scale compute infrastructure and utilizing operational data to enhance availability, performance, and fleet efficiency.
β’ In-depth knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.
β’ Understanding of failure modes associated with large distributed AI workloads and the necessary infrastructure features to consistently support prolonged training and production inference.
β’ Equity
β’ Benefits
TMS
StaffMill Oy
HumanIT Digital Consulting
HumanIT Digital Consulting
Get handpicked remote jobs straight to your inbox weekly.