Senior Engineer

Posted Aug 24

This is a fully remote position, open to applicants in California, +1 more state.

πŸ“‹ Description

β€’ Oversee NVIDIA Cloud Partner Day 2 operational readiness initiatives post-initial deployment and activation.

β€’ Partner with NVIDIA Cloud Partners to create systems, procedures, automation, and operational methodologies for optimized infrastructure.

β€’ Develop ongoing validation for GPU, CPU, storage, and network health within extensive AI clusters.

β€’ Set up telemetry, monitoring, alerting, dashboards, and operational indicators across compute, GPU, networking, storage, Kubernetes, and AI workloads.

β€’ Create automated workflows to detect, isolate, drain, repair, validate, and restore non-functional infrastructure back to service.

β€’ Oversee the GPU fleet lifecycle, encompassing drivers, firmware, Kubernetes nodes, OS updates, configuration management, upgrades, and configuration drift.

β€’ Convert NVIDIA NCP requirements and reference architectures into production operational practices, validation standards, runbooks, automation, and measurable benchmarks.

β€’ Define health indicators, SLOs, metrics, acceptance standards, and infrastructure readiness validation.

β€’ Develop reusable tools, automation, implementation guides, runbooks, playbooks, and reference implementations throughout NCP environments.


⛳️ Requirements

β€’ BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent experience.

β€’ Over 8 years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or comparable roles in large-scale production settings.

β€’ Extensive experience in managing Linux-based distributed systems and cloud infrastructure in a production environment.

β€’ Profound knowledge of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.

β€’ Strong grasp of production observability, including metrics, logging, alerting, dashboards, health checks, and operations driven by service level agreements.

β€’ Experience in creating automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.

β€’ Solid networking principles and experience troubleshooting complex distributed systems across compute, network, and storage layers.

β€’ Programming and automation skills using Python, Go, shell scripting, or similar languages.

β€’ Experience managing large GPU or accelerated computing infrastructure that supports AI training and inference workloads.

β€’ Familiarity with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.

β€’ Proven track record of collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure while operating SLOs for large-scale compute infrastructure and utilizing operational data to enhance availability, performance, and fleet efficiency.

β€’ In-depth knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.

β€’ Understanding of failure modes associated with large distributed AI workloads and the necessary infrastructure features to consistently support prolonged training and production inference.


🏝️ Benefits

β€’ Equity

β€’ Benefits

People also viewed

TMS7 hours ago

.Net Full Stack Lead, Kentico CMS, Genesys Cloud

US flagCalifornia OnlyFreelanceFull-stack Engineer
ApplyView job
StaffMill Oy7 hours ago

Junior Software Developer

FI flagFinland OnlyPart-timeFull-stack Engineer
ApplyView job
HumanIT Digital Consulting7 hours ago

Senior/Staff Java Fullstack Developer – Java, Spring Boot, React

PT flagPortugal OnlyFull-timeFull-stack Engineer
ApplyView job
HumanIT Digital Consulting7 hours ago

Fullstack Developer, Java, React, AI Tooling

PT flagPortugal OnlyFull-timeFull-stack Engineer€2,550 – €3,350/month
ApplyView job
Humana7 hours ago

Manager, Full Stack Engineering

US flagFlorida, +8 more statesFull-timeFull-stack Engineer$117.6k – $161.7k/year
ApplyView job
Capital Technology Group, LLC7 hours ago

Software Developer

US flagUnited States OnlyFull-timeFull-stack Engineer$75k – $110k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers