
Senior Solutions Architect – AI Infrastructure Networking
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
• Take ownership of network solutions for k0rdent AI, Mirantis' platform dedicated to constructing and managing GPU clouds and AI factories.
• Create designs for large GPU cluster interconnections, tenant isolation, and workload connectivity across containers, virtual machines, and bare-metal nodes.
• Develop and publish network reference architectures and solution designs ranging from single-rack setups to multi-thousand-GPU clusters.
• Define InfiniBand and RoCEv2 Ethernet compute fabrics, including rail-optimized and fat-tree/Clos designs, along with oversubscription and failure domains.
• Outline front-end, storage, and management networks, ensuring connectivity to customer data centers, public clouds, and hybrid environments.
• Specify multi-tenant isolation leveraging InfiniBand PKeys, VRFs, EVPN-VXLAN, and Kubernetes network policies.
• Document scalability limits and trade-offs related to cost, performance, operability, and vendor lock-in.
• Define Linux host networking for NICs, DPUs, and SuperNICs, incorporating PCI passthrough, SR-IOV, IOMMU, NUMA, and GPU-NIC affinity.
• Design Kubernetes networking tailored for AI workloads utilizing CNIs, Multus, SR-IOV, RDMA plugins, NVIDIA Network Operator, and DRA.
• Create networking solutions for KubeVirt concerning GPU and RDMA traffic.
• Investigate emerging AI networking technologies, including Ultra Ethernet, SuperNIC/DPU offloads, scale-up/scale-across fabrics, and DRA network drivers.
• Prototype and benchmark various designs in a laboratory environment.
• Publish research notes, design proposals, papers, blog posts, and presentations.
• Construct and manage customer and partner proofs of concept on Mirantis lab hardware and at customer locations.
• Develop automation and tooling using Python, Go, Bash, Ansible, Helm, Kubernetes manifests, and Terraform.
• Validate and benchmark fabrics and host configurations through NCCL tests, perftest, and ib_write_bw.
• Serve as a network subject matter expert during customer discovery, design reviews, and architecture workshops.
• Collaborate with hardware and networking partners on joint designs and validations.
• Integrate research findings into Product and Engineering to help shape the k0rdent AI roadmap.
• Present at industry events, webinars, and partner summits.
• Conduct technical workshops and create reference architecture documents, solution briefs, blog posts, and internal training materials.
• A Bachelor's degree in Computer Science, Electrical Engineering, Telecommunications, or a related field, or equivalent practical experience.
• Over 8 years of experience in network engineering or network architecture, including a minimum of 3 years in data center, HPC, or cloud infrastructure networking.
• Experience in customer-facing roles as a solutions architect, pre-sales engineer, consultant, or technical lead.
• Proficiency in data center network design, including spine-leaf and Clos topologies, rail-optimized GPU fabrics, oversubscription, and ECMP.
• Expertise in InfiniBand subnet management, partitioning, adaptive routing, and NCCL/RDMA traffic.
• Knowledge of RoCEv2, PFC, ECN, DCQCN, QoS, MTU, and buffer tuning.
• Familiarity with BGP, EVPN-VXLAN, VRFs, hybrid and multi-cloud connectivity, VPNs, IP address, and DNS planning.
• Understanding of Linux networking, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool, and devlink.
• Knowledge of Kubernetes networking, CNI plugins, Multus, SR-IOV, RDMA device plugins, network policy, and service exposure.
• Experience with KubeVirt VM networking, passthrough, and SR-IOV.
• Proficient in programming or scripting in at least one language; Python or Go preferred.
• Comfortable using Git, CI, and infrastructure-as-code practices.
• Excellent written and spoken English skills.
• Capable of presenting to large audiences and conducting hands-on workshops.
• Experience working in an international, distributed company across various time zones and cultures.
• Willingness to participate in meetings outside of standard local hours.
• Must reside in Europe.
• Additional qualifications may include expertise in NVIDIA networking, GPU cloud/HPC, bare-metal provisioning, network automation, SDN, load balancing, storage networking, open-source contributions, standards participation, and proficiency in additional European languages.
• Option to work remotely within Europe.
• Travel opportunities of up to 25% for customer engagements, partner meetings, lab work, and industry events.
• Compensation and benefits aligned with local market standards and employment arrangements.
• Collaborate with a distributed, international team.
• Opportunities to influence product narratives and contribute to go-to-market strategies.
• Direct engagement with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.
FineCom Logistics
Mirantis
Omtera
Pearson VUE
Get handpicked remote jobs straight to your inbox weekly.