Remotery

Senior Solutions Architect, Cloud Partner Operations

atNVIDIARemoteUS flagCaliforniaFull-timeSolutions EngineerSenior$224k – $356.5k/year

Posted 2 days ago

This is a fully remote position, open to applicants in California.

📋 Description

• Tackle challenging Day 2 operations issues at scale in collaboration with partner engineers.

• Investigate root causes, prototype solutions, validate them under realistic loads, and equip partners with actionable practices.

• Prepare partners for new NVIDIA platforms, including capacity, services, and use cases.

• Facilitate adoption in live settings while maintaining service quality.

• Enhance reliability, performance, utilization, recovery time, and cost efficiency per token.

• Identify and assist in bridging maturity gaps in people, processes, tools, telemetry, security, and incident response.

• Transform validated efforts into operational procedures, reference architectures, assessments, automation, and agentic workflows.

• Recognize cross-partner trends and provide field data to account teams, support, product, and engineering.

• Enhance NVIDIA Cloud Partner Day 2 operations and ecosystem capabilities.


⛳️ Requirements

• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.

• 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of distinguished specialist-level experience in large-scale GPU or AI infrastructure.

• Proven experience in building, operating, or enhancing distributed infrastructure under actual production load.

• In-depth knowledge in at least one aspect of the Day 2 stack, with practical large-scale GPU, HPC, or cloud infrastructure experience.

• Familiarity with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or similar platforms.

• Hands-on experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or comparable tools.

• Strong knowledge of Linux operating systems.

• Proficiency in Python, Bash, or similar scripting languages for automation purposes.

• Evidence-based troubleshooting capabilities across system boundaries.

• Ability to lead complex initiatives with partner engineers and cross-functional teams without direct authority.

• Excellent communication, prioritization, and time-management abilities across multiple partner interactions.

• Preferred/standout experience in operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; experience in establishing 24/7 operations; familiarity with NVIDIA rack-scale platforms; knowledge of NVIDIA operations technologies; and improvements in fleet health or unit economics.


🏝️ Benefits

• Competitive salaries

• Generous benefits package

• Equity

People also viewed

Whippy11 hours ago

Solutions Engineer

US flagUnited States OnlyFull-timeSolutions Engineer
ApplyView job
Snowflake13 hours ago

Senior Solution Engineer

US flagFlorida, +3 more statesFull-timeSolutions Engineer$138k – $181.1k/year
ApplyView job
Cisco13 hours ago

Partner Solutions Engineer

US flagNew Jersey OnlyFull-timeSolutions Engineer$212.2k – $268.1k/year
ApplyView job
BCD Travel15 hours ago

AI Solutions Engineer – Architect

GB flagUnited Kingdom, +2 more statesFull-timeSolutions Engineer€110k – €190k/year
ApplyView job
AutoStore™17 hours ago

Solutions Consultant – Warehouse Automation

AU flagAustralia OnlyFull-timeSolutions Engineer
ApplyView job
Agility Technologies Inc23 hours ago

Solutions Architect, Databricks

US flagUnited States OnlyFull-timeSolutions Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers