
Senior Solutions Architect – AI Infrastructure Networking
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States, +1 more country.
• Take ownership of the network solutions for k0rdent AI, Mirantis's platform dedicated to building and operating GPU clouds and AI factories.
• Create interconnects for extensive GPU clusters, ensuring tenant isolation and workload connectivity across containers, virtual machines, and bare-metal nodes.
• Design and publish network reference architectures and solution designs that range from single-rack setups to multi-thousand-GPU clusters.
• Define compute fabrics for InfiniBand and RoCEv2 Ethernet, including front-end, storage, and management networks.
• Specify multi-tenant isolation using InfiniBand PKeys, VRFs, EVPN-VXLAN, and Kubernetes network policies.
• Document scalability limits and trade-offs regarding cost, performance, operability, and vendor lock-in.
• Define Linux host networking involving NICs, DPUs, SuperNICs, PCI passthrough, SR-IOV, IOMMU, NUMA, and GPU-NIC affinity.
• Design networking for AI workloads in Kubernetes and KubeVirt, incorporating CNI, Multus, RDMA, SR-IOV, and DRA.
• Investigate and assess emerging AI networking technologies and standards.
• Prototype and evaluate alternative designs in both lab environments and with customers.
• Publish internal research notes, design proposals, papers, blog posts, and presentations.
• Develop proofs of concept for customers and partners using Mirantis and customer hardware.
• Write automation scripts and tools using Python, Go, Bash, Ansible, Helm, Kubernetes manifests, and Terraform.
• Validate and benchmark fabrics and host configurations utilizing tools such as NCCL tests, perftest, and ib_write_bw.
• Act as a network subject matter expert during customer discovery sessions, design reviews, and architecture workshops.
• Collaborate with teams across hardware, networking, Solutions Architecture, Partner Management, and Engineering.
• Contribute research findings to Product and Engineering, assisting in shaping the k0rdent AI roadmap.
• Present at industry events, webinars, and partner summits, and conduct technical workshops.
• Produce reference architectures, solution briefs, blog posts, and internal enablement materials.
• Bachelor’s degree in Computer Science, Electrical Engineering, Telecommunications, or a related field, or equivalent practical experience.
• Over 8 years of experience in network engineering or network architecture.
• Minimum of 3 years in networking for data centers, HPC, or cloud infrastructure.
• Customer-facing experience in roles such as solutions architect, pre-sales engineer, consultant, or technical lead.
• Expertise in data center network design, including spine-leaf and Clos topologies, rail-optimized GPU fabrics, oversubscription, ECMP, and scaling to thousands of nodes.
• Knowledge of InfiniBand subnet management, partitioning, adaptive routing, and NCCL/RDMA traffic.
• Familiarity with RoCEv2 Ethernet technologies: PFC, ECN, DCQCN, QoS, MTU, and buffer tuning.
• Experience with BGP, EVPN-VXLAN, and VRFs for multi-tenancy.
• Knowledge of hybrid and multi-cloud connectivity, interconnects, VPN, cloud networking, IP address, and DNS planning.
• Proficiency in Linux networking, including NIC drivers, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool, and devlink.
• Expertise in Kubernetes networking, including CNI plugins, Multus, SR-IOV, and RDMA device plugins, as well as network policy and service exposure.
• Knowledge of KubeVirt VM networking, including passthrough and SR-IOV.
• Programming or scripting experience in at least one language, preferably Python or Go.
• Comfortable using Git, CI, and infrastructure-as-code methodologies.
• Excellent proficiency in written and spoken English.
• Comfortable presenting at conferences and events and leading hands-on workshops.
• Experience working in an international and distributed company across diverse time zones and work cultures.
• This role is open to candidates located in Europe, the United States (East Coast preferred), or India.
• Availability for some meetings outside of standard local hours.
• Preferred experience includes NVIDIA networking, GPU cloud/HPC, bare-metal provisioning, network automation, SDN controllers, load balancing, NVMe-oF, Mirantis products, open-source contributions, standards participation, published research, and proficiency in additional languages.
• Remote work arrangement.
• Travel up to 25% for customer engagements, partner meetings, lab work, and industry events.
• Collaboration within a distributed, international team.
• Opportunity to work directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.
• Chance to influence product narrative and contribute to go-to-market success.
• Compensation and benefits tailored to local market conditions and employment arrangements.
Abnormal Security
Everbridge
Centene Corporation
Get handpicked remote jobs straight to your inbox weekly.