Senior Cloud Infrastructure and Network Operations Solutions Architect

Posted 4 days ago

This is a fully remote position, open to applicants in United Kingdom, +3 more countries.

📋 Description

• Collaborate directly with customers, partners, and cross-functional teams to evaluate, design, and manage compute, storage, and scale-up fabrics that support extensive GPU deployments.

• Take ownership of the fabric aspect of NVIDIA’s Cloud Partner operating model throughout the entire Day 1 to Day 2 lifecycle.

• Oversee NVLink/NVSwitch partition operations, ensuring maintenance-partition isolation and secure partition-change workflows for multi-tenant environments.

• Manage the lifecycle of switch software and firmware, which includes Cumulus Linux, SONiC, and switch-OS upgrades, as well as the execution of CPLD and field-notice rollout campaigns.

• Assist in Day 1 fabric validation and acceptance, encompassing InfiniBand/UFM bring-up, Spectrum-X/RoCE Ethernet setup, cabling and link-health verification, routing and congestion-control validation, along with multi-day burn-in processes.

• Reduce the time from cluster handover to the initiation of the first production workload by eliminating redundant fabric validation.

• Enhance fabric reliability at fleet scale through telemetry, fault detection, remediation, root-cause analysis, and improvement of MTBI and job goodput.

• Offer consultative support and hands-on troubleshooting across NICs, DPUs, switch OS, routing, congestion control, host networking, and Kubernetes integration.

• Serve as a technical leader for designated accounts, conducting structured knowledge transfer and creating runbooks for partner teams.


⛳️ Requirements

• BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.

• Over 5 years of professional experience in data center networking, fabric engineering, or large-scale network operations positions.

• Profound understanding of data center network architectures and RDMA fabrics—InfiniBand and RoCE/Ethernet—covering topology and routing design, congestion control, lossless/QoS configuration, and troubleshooting across NICs, switches, and high-speed interconnects.

• Practical experience with InfiniBand and UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX and BlueField NICs/DPUs, and NVLink/NVSwitch systems.

• Extensive knowledge of Linux (RedHat, Ubuntu), switch operating systems, network OS internals, host networking, and OS-level security.

• Familiarity with storage and compute traffic patterns in HPC/AI clusters.

• Proficiency in Python and Bash scripting, configuration management, Infrastructure-as-Code tools like Ansible and Terraform, GitOps-based network configuration and firmware/upgrade management, and observability stacks such as Grafana, Loki, and Prometheus.

• Proven capability to measure and enhance MTBI and job goodput on large GPU clusters, including fault detection, drain and remediation workflows, SLO/error-budget definition, and post-incident review.

• Strong consultative experience leading architectural reviews and presenting to executive stakeholders.

• Experience with Kubernetes networking in GPU clusters, including CNI plugins, multi-network attachment, SR-IOV, and RDMA device plugins.

• Familiarity with NVIDIA Network and GPU Operators.

• Expertise in DOCA and DPU infrastructure services such as DOCA DPF.

• Knowledge of fabric and GPU health telemetry, including DCGM and XID diagnostics, link-level and switch counters, node-level health agents, and fleet-wide reliability intelligence.

• Understanding of large-scale training traffic behavior, encompassing NCCL tuning, rail alignment, and diagnosing network-bound performance regressions.

• Experience in delivering multi-tenant network isolation using VRF/VLAN/EVPN segmentation, tenant partitioning, and secure fabric handover.


🏝️ Benefits

• Competitive salary and performance-based incentives.

• Comprehensive health and wellness benefits.

• Opportunities for professional development and continuous learning.

• Collaborative and inclusive company culture.

• Flexible working hours and remote work options.

People also viewed

Signal & Strand1 day ago

Principal Time & Attendance Solutions Consultant

US flagCalifornia, +3 more statesFull-timeSolutions Engineer$180k – $210k/year
ApplyView job
Perseus Group, Constellation Software1 day ago

Sales Solutions Consultant – Account Management

US flagFlorida OnlyFull-timeSolutions Engineer$108k – $132k/year
ApplyView job
GP Software1 day ago

Technology Solutions Engineer

FR flagFrance OnlyFull-timeSolutions Engineer€24k – €36k/year
ApplyView job
McKesson1 day ago

Lead Intelligent Solutions Engineer

US flagTexas OnlyFull-timeSolutions Engineer
ApplyView job
McKesson1 day ago

Lead Intelligent Solutions Engineer

US flagTexas OnlyFull-timeSolutions Engineer
ApplyView job
Teleperformance1 day ago

Conversational AI Solution Engineer

BA flagBosnia and Herzegovina OnlyFull-timeSolutions Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers