
Senior Cloud Infrastructure and Network Operations Solutions Architect
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United Kingdom, +3 more countries.
• Collaborate directly with customers, partners, and cross-functional teams to evaluate, design, and manage compute, storage, and scale-up fabrics that support extensive GPU deployments.
• Take ownership of the fabric aspect of NVIDIA’s Cloud Partner operating model throughout the entire Day 1 to Day 2 lifecycle.
• Oversee NVLink/NVSwitch partition operations, ensuring maintenance-partition isolation and secure partition-change workflows for multi-tenant environments.
• Manage the lifecycle of switch software and firmware, which includes Cumulus Linux, SONiC, and switch-OS upgrades, as well as the execution of CPLD and field-notice rollout campaigns.
• Assist in Day 1 fabric validation and acceptance, encompassing InfiniBand/UFM bring-up, Spectrum-X/RoCE Ethernet setup, cabling and link-health verification, routing and congestion-control validation, along with multi-day burn-in processes.
• Reduce the time from cluster handover to the initiation of the first production workload by eliminating redundant fabric validation.
• Enhance fabric reliability at fleet scale through telemetry, fault detection, remediation, root-cause analysis, and improvement of MTBI and job goodput.
• Offer consultative support and hands-on troubleshooting across NICs, DPUs, switch OS, routing, congestion control, host networking, and Kubernetes integration.
• Serve as a technical leader for designated accounts, conducting structured knowledge transfer and creating runbooks for partner teams.
• BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
• Over 5 years of professional experience in data center networking, fabric engineering, or large-scale network operations positions.
• Profound understanding of data center network architectures and RDMA fabrics—InfiniBand and RoCE/Ethernet—covering topology and routing design, congestion control, lossless/QoS configuration, and troubleshooting across NICs, switches, and high-speed interconnects.
• Practical experience with InfiniBand and UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX and BlueField NICs/DPUs, and NVLink/NVSwitch systems.
• Extensive knowledge of Linux (RedHat, Ubuntu), switch operating systems, network OS internals, host networking, and OS-level security.
• Familiarity with storage and compute traffic patterns in HPC/AI clusters.
• Proficiency in Python and Bash scripting, configuration management, Infrastructure-as-Code tools like Ansible and Terraform, GitOps-based network configuration and firmware/upgrade management, and observability stacks such as Grafana, Loki, and Prometheus.
• Proven capability to measure and enhance MTBI and job goodput on large GPU clusters, including fault detection, drain and remediation workflows, SLO/error-budget definition, and post-incident review.
• Strong consultative experience leading architectural reviews and presenting to executive stakeholders.
• Experience with Kubernetes networking in GPU clusters, including CNI plugins, multi-network attachment, SR-IOV, and RDMA device plugins.
• Familiarity with NVIDIA Network and GPU Operators.
• Expertise in DOCA and DPU infrastructure services such as DOCA DPF.
• Knowledge of fabric and GPU health telemetry, including DCGM and XID diagnostics, link-level and switch counters, node-level health agents, and fleet-wide reliability intelligence.
• Understanding of large-scale training traffic behavior, encompassing NCCL tuning, rail alignment, and diagnosing network-bound performance regressions.
• Experience in delivering multi-tenant network isolation using VRF/VLAN/EVPN segmentation, tenant partitioning, and secure fabric handover.
• Competitive salary and performance-based incentives.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive company culture.
• Flexible working hours and remote work options.
Signal & Strand
Perseus Group, Constellation Software
GP Software
McKesson
Get handpicked remote jobs straight to your inbox weekly.