
Senior Solutions Architect, Cloud Partner Operations
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Tackle challenging Day 2 operations issues at scale in collaboration with partner engineers.
• Investigate root causes, prototype solutions, validate them under realistic loads, and equip partners with actionable practices.
• Prepare partners for new NVIDIA platforms, including capacity, services, and use cases.
• Facilitate adoption in live settings while maintaining service quality.
• Enhance reliability, performance, utilization, recovery time, and cost efficiency per token.
• Identify and assist in bridging maturity gaps in people, processes, tools, telemetry, security, and incident response.
• Transform validated efforts into operational procedures, reference architectures, assessments, automation, and agentic workflows.
• Recognize cross-partner trends and provide field data to account teams, support, product, and engineering.
• Enhance NVIDIA Cloud Partner Day 2 operations and ecosystem capabilities.
• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.
• 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of distinguished specialist-level experience in large-scale GPU or AI infrastructure.
• Proven experience in building, operating, or enhancing distributed infrastructure under actual production load.
• In-depth knowledge in at least one aspect of the Day 2 stack, with practical large-scale GPU, HPC, or cloud infrastructure experience.
• Familiarity with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or similar platforms.
• Hands-on experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or comparable tools.
• Strong knowledge of Linux operating systems.
• Proficiency in Python, Bash, or similar scripting languages for automation purposes.
• Evidence-based troubleshooting capabilities across system boundaries.
• Ability to lead complex initiatives with partner engineers and cross-functional teams without direct authority.
• Excellent communication, prioritization, and time-management abilities across multiple partner interactions.
• Preferred/standout experience in operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; experience in establishing 24/7 operations; familiarity with NVIDIA rack-scale platforms; knowledge of NVIDIA operations technologies; and improvements in fleet health or unit economics.
• Competitive salaries
• Generous benefits package
• Equity
Snowflake
Cisco
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.