Senior Solutions Architect – AI Infrastructure Networking

Posted 2 days ago

This is a fully remote position, open to applicants in Europe.

📋 Description

• Take ownership of network solutions for k0rdent AI, Mirantis' platform dedicated to constructing and managing GPU clouds and AI factories.

• Create designs for large GPU cluster interconnections, tenant isolation, and workload connectivity across containers, virtual machines, and bare-metal nodes.

• Develop and publish network reference architectures and solution designs ranging from single-rack setups to multi-thousand-GPU clusters.

• Define InfiniBand and RoCEv2 Ethernet compute fabrics, including rail-optimized and fat-tree/Clos designs, along with oversubscription and failure domains.

• Outline front-end, storage, and management networks, ensuring connectivity to customer data centers, public clouds, and hybrid environments.

• Specify multi-tenant isolation leveraging InfiniBand PKeys, VRFs, EVPN-VXLAN, and Kubernetes network policies.

• Document scalability limits and trade-offs related to cost, performance, operability, and vendor lock-in.

• Define Linux host networking for NICs, DPUs, and SuperNICs, incorporating PCI passthrough, SR-IOV, IOMMU, NUMA, and GPU-NIC affinity.

• Design Kubernetes networking tailored for AI workloads utilizing CNIs, Multus, SR-IOV, RDMA plugins, NVIDIA Network Operator, and DRA.

• Create networking solutions for KubeVirt concerning GPU and RDMA traffic.

• Investigate emerging AI networking technologies, including Ultra Ethernet, SuperNIC/DPU offloads, scale-up/scale-across fabrics, and DRA network drivers.

• Prototype and benchmark various designs in a laboratory environment.

• Publish research notes, design proposals, papers, blog posts, and presentations.

• Construct and manage customer and partner proofs of concept on Mirantis lab hardware and at customer locations.

• Develop automation and tooling using Python, Go, Bash, Ansible, Helm, Kubernetes manifests, and Terraform.

• Validate and benchmark fabrics and host configurations through NCCL tests, perftest, and ib_write_bw.

• Serve as a network subject matter expert during customer discovery, design reviews, and architecture workshops.

• Collaborate with hardware and networking partners on joint designs and validations.

• Integrate research findings into Product and Engineering to help shape the k0rdent AI roadmap.

• Present at industry events, webinars, and partner summits.

• Conduct technical workshops and create reference architecture documents, solution briefs, blog posts, and internal training materials.


⛳️ Requirements

• A Bachelor's degree in Computer Science, Electrical Engineering, Telecommunications, or a related field, or equivalent practical experience.

• Over 8 years of experience in network engineering or network architecture, including a minimum of 3 years in data center, HPC, or cloud infrastructure networking.

• Experience in customer-facing roles as a solutions architect, pre-sales engineer, consultant, or technical lead.

• Proficiency in data center network design, including spine-leaf and Clos topologies, rail-optimized GPU fabrics, oversubscription, and ECMP.

• Expertise in InfiniBand subnet management, partitioning, adaptive routing, and NCCL/RDMA traffic.

• Knowledge of RoCEv2, PFC, ECN, DCQCN, QoS, MTU, and buffer tuning.

• Familiarity with BGP, EVPN-VXLAN, VRFs, hybrid and multi-cloud connectivity, VPNs, IP address, and DNS planning.

• Understanding of Linux networking, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool, and devlink.

• Knowledge of Kubernetes networking, CNI plugins, Multus, SR-IOV, RDMA device plugins, network policy, and service exposure.

• Experience with KubeVirt VM networking, passthrough, and SR-IOV.

• Proficient in programming or scripting in at least one language; Python or Go preferred.

• Comfortable using Git, CI, and infrastructure-as-code practices.

• Excellent written and spoken English skills.

• Capable of presenting to large audiences and conducting hands-on workshops.

• Experience working in an international, distributed company across various time zones and cultures.

• Willingness to participate in meetings outside of standard local hours.

• Must reside in Europe.

• Additional qualifications may include expertise in NVIDIA networking, GPU cloud/HPC, bare-metal provisioning, network automation, SDN, load balancing, storage networking, open-source contributions, standards participation, and proficiency in additional European languages.


🏝️ Benefits

• Option to work remotely within Europe.

• Travel opportunities of up to 25% for customer engagements, partner meetings, lab work, and industry events.

• Compensation and benefits aligned with local market standards and employment arrangements.

• Collaborate with a distributed, international team.

• Opportunities to influence product narratives and contribute to go-to-market strategies.

• Direct engagement with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.

People also viewed

FineCom Logistics18 hours ago

Integration Engineer

DE flagGermany OnlyFull-timeSolutions Engineer
ApplyView job
Mirantis22 hours ago

Senior Solutions Architect – AI Infrastructure Security

EuropeFull-timeSolutions Engineer
ApplyView job
Omtera22 hours ago

Solutions Consultant, Product Analytics & Growth

SA flagSaudi Arabia OnlyFull-timeSolutions Engineer
ApplyView job
Pearson VUE1 day ago

Lead Specialist, Solution Architect

US flagUnited States OnlyFull-timeSolutions Engineer$150k – $160k/year
ApplyView job
The Hanover Insurance Group1 day ago

Personal Lines Risk Solutions Consultant

US flagConnecticut, +1 more stateFull-timeSolutions Engineer
ApplyView job
Redtech1 day ago

Solutions Architect, Microsoft Power Platform, Azure

US flagWashington OnlyFreelanceSolutions Engineer$70 – $80/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers