
Senior Solutions Architect, Cloud Partner Operations
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in California.
• Address complex Day 2 operations challenges at scale in collaboration with partner engineers.
• Investigate root causes, prototype solutions, validate findings under realistic conditions, and establish operational practices for partners to maintain.
• Assist partners in developing operational models for new NVIDIA platforms, capacity, services, and use cases.
• Facilitate adoption in live settings without compromising service quality.
• Enhance reliability, performance, and cost-effectiveness by utilizing metrics such as incident frequency, recovery time, utilization, and cost-per-token.
• Identify and assist in bridging Day 2 maturity gaps across personnel, processes, tools, telemetry, security, and incident response.
• Transform validated findings into operational procedures, reference architectures, assessments, automation, and agentic workflows.
• Recognize cross-partner trends and provide field data to account teams, support, product, and engineering.
• Enhance NVIDIA's factory planning capabilities.
• BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related discipline - or equivalent experience.
• 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a related technical position; alternatively, 5+ years of outstanding specialist-level experience in large-scale GPU or AI infrastructure.
• Proven experience in building, operating, or enhancing distributed infrastructure under actual production conditions.
• Extensive knowledge in at least one component of the Day 2 stack, supported by hands-on experience with large-scale GPU, HPC, or cloud infrastructure.
• Practical experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation tools like Terraform, Ansible, Argo CD, or similar.
• Strong Linux proficiency and sufficient experience in Python, Bash, or similar for automating measurement, diagnosis, validation, or remediation tasks.
• Comprehensive evidence-based troubleshooting across system boundaries.
• Capability to lead intricate projects with partner engineers and cross-functional teams without direct authority.
• Excellent communication, prioritization, and time-management skills across multiple partner collaborations.
• Practical experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer demand.
• Background in establishing or advancing a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.
• Hands-on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72, or NVIDIA operations technologies like Spectrum-X, UFM, Base Command Manager, Mission Control, and GPU or Network Operators.
• Experience in enhancing fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnostics, or agent-based remediation.
• Competitive salaries.
• Generous benefits package.
• Equity.
• Additional benefits.
Databricks
Indigo Beam Consulting
MCA Connect
Cisco
Get handpicked remote jobs straight to your inbox weekly.