Senior Solutions Architect – AI Infrastructure Networking

Posted 2 days ago

This is a fully remote position, open to applicants in United States, +1 more country.

📋 Description

• Take ownership of the network solutions for k0rdent AI, Mirantis's platform dedicated to building and operating GPU clouds and AI factories.

• Create interconnects for extensive GPU clusters, ensuring tenant isolation and workload connectivity across containers, virtual machines, and bare-metal nodes.

• Design and publish network reference architectures and solution designs that range from single-rack setups to multi-thousand-GPU clusters.

• Define compute fabrics for InfiniBand and RoCEv2 Ethernet, including front-end, storage, and management networks.

• Specify multi-tenant isolation using InfiniBand PKeys, VRFs, EVPN-VXLAN, and Kubernetes network policies.

• Document scalability limits and trade-offs regarding cost, performance, operability, and vendor lock-in.

• Define Linux host networking involving NICs, DPUs, SuperNICs, PCI passthrough, SR-IOV, IOMMU, NUMA, and GPU-NIC affinity.

• Design networking for AI workloads in Kubernetes and KubeVirt, incorporating CNI, Multus, RDMA, SR-IOV, and DRA.

• Investigate and assess emerging AI networking technologies and standards.

• Prototype and evaluate alternative designs in both lab environments and with customers.

• Publish internal research notes, design proposals, papers, blog posts, and presentations.

• Develop proofs of concept for customers and partners using Mirantis and customer hardware.

• Write automation scripts and tools using Python, Go, Bash, Ansible, Helm, Kubernetes manifests, and Terraform.

• Validate and benchmark fabrics and host configurations utilizing tools such as NCCL tests, perftest, and ib_write_bw.

• Act as a network subject matter expert during customer discovery sessions, design reviews, and architecture workshops.

• Collaborate with teams across hardware, networking, Solutions Architecture, Partner Management, and Engineering.

• Contribute research findings to Product and Engineering, assisting in shaping the k0rdent AI roadmap.

• Present at industry events, webinars, and partner summits, and conduct technical workshops.

• Produce reference architectures, solution briefs, blog posts, and internal enablement materials.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Electrical Engineering, Telecommunications, or a related field, or equivalent practical experience.

• Over 8 years of experience in network engineering or network architecture.

• Minimum of 3 years in networking for data centers, HPC, or cloud infrastructure.

• Customer-facing experience in roles such as solutions architect, pre-sales engineer, consultant, or technical lead.

• Expertise in data center network design, including spine-leaf and Clos topologies, rail-optimized GPU fabrics, oversubscription, ECMP, and scaling to thousands of nodes.

• Knowledge of InfiniBand subnet management, partitioning, adaptive routing, and NCCL/RDMA traffic.

• Familiarity with RoCEv2 Ethernet technologies: PFC, ECN, DCQCN, QoS, MTU, and buffer tuning.

• Experience with BGP, EVPN-VXLAN, and VRFs for multi-tenancy.

• Knowledge of hybrid and multi-cloud connectivity, interconnects, VPN, cloud networking, IP address, and DNS planning.

• Proficiency in Linux networking, including NIC drivers, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool, and devlink.

• Expertise in Kubernetes networking, including CNI plugins, Multus, SR-IOV, and RDMA device plugins, as well as network policy and service exposure.

• Knowledge of KubeVirt VM networking, including passthrough and SR-IOV.

• Programming or scripting experience in at least one language, preferably Python or Go.

• Comfortable using Git, CI, and infrastructure-as-code methodologies.

• Excellent proficiency in written and spoken English.

• Comfortable presenting at conferences and events and leading hands-on workshops.

• Experience working in an international and distributed company across diverse time zones and work cultures.

• This role is open to candidates located in Europe, the United States (East Coast preferred), or India.

• Availability for some meetings outside of standard local hours.

• Preferred experience includes NVIDIA networking, GPU cloud/HPC, bare-metal provisioning, network automation, SDN controllers, load balancing, NVMe-oF, Mirantis products, open-source contributions, standards participation, published research, and proficiency in additional languages.


🏝️ Benefits

• Remote work arrangement.

• Travel up to 25% for customer engagements, partner meetings, lab work, and industry events.

• Collaboration within a distributed, international team.

• Opportunity to work directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.

• Chance to influence product narrative and contribute to go-to-market success.

• Compensation and benefits tailored to local market conditions and employment arrangements.

People also viewed

Abnormal Security1 day ago

Customer Solutions Engineer

GB flagUnited Kingdom OnlyFull-timeSolutions Engineer
ApplyView job
Darede1 day ago

Mid-Level Solutions Architect

BR flagBrazil OnlyFull-timeSolutions Engineer
ApplyView job
Everbridge1 day ago

Solutions Consultant

US flagUnited States OnlyFull-timeSolutions Engineer$94k – $130k/year
ApplyView job
Centene Corporation1 day ago

Lead Business Solutions Developer

US flagMissouri OnlyFull-timeSolutions Engineer$107.7k – $199.3k/year
ApplyView job
Mirantis1 day ago

Senior Solutions Architect – AI Infrastructure Security

US flagMassachusetts OnlyFull-timeSolutions Engineer
ApplyView job
Palo Alto Networks1 day ago

Solutions Consultant 2 – Healthcare

US flagFlorida OnlyFull-timeSolutions Engineer$206.4k – $283.8k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers