
Senior Software Engineer β Cluster Networking
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California, +4 more states.
β’ Take ownership of and advance the Kubernetes networking framework for GPU clusters operating at a scale of several thousand nodes.
β’ Architect, manage, and enhance the overlay network, including CNI, mesh, and VPN topologies (Tailscale, WireGuard), as well as the gateways that link control and data planes.
β’ Design, manage, and enhance L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
β’ Identify and resolve scalability limitations such as packet loss under heavy load, control-plane saturation, IP address management exhaustion, and failure modes when exceeding a few thousand nodes.
β’ Develop scale-testing environments and validation frameworks to detect networking regressions prior to production deployment.
β’ Troubleshoot complex and ambiguous issues throughout the stack, tracing symptoms in Slurm or training jobs back to their networking root causes.
β’ Collaborate with cloud and neocloud providers to define network topology, requirements, and capabilities for new clusters.
β’ Offer senior-level technical expertise to a distributed team within the Cluster Networking domain.
β’ BS/MS in Computer Science, Electrical Engineering, or a related discipline, or equivalent professional experience.
β’ Over 6 years of professional experience in systems, networking, or infrastructure software engineering.
β’ Extensive knowledge of Kubernetes networking architecture and CNI standards, with hands-on experience operating Calico highly preferred.
β’ Expertise in designing and maintaining contemporary mesh and VPN networking topologies such as Tailscale, WireGuard, or similar solutions.
β’ Strong foundational knowledge of Linux networking: routing, netfilter and iptables/nftables, packet marking, network namespaces, and their interaction with container runtimes.
β’ Proven ability to debug distributed network issues at scale, including packet capture, tracing, and correlating behavior across numerous hosts to identify a single root cause.
β’ Proficiency in Go, Python, C, or a similar systems programming language.
β’ Excellent written and verbal communication skills, with the capacity to collaborate effectively with engineers across various time zones.
β’ Direct experience in architecting and managing large-scale Kubernetes configurations involving thousands of concurrent nodes.
β’ Familiarity with high-performance fabrics in AI or HPC settings, including InfiniBand, RoCE, or RDMA over converged networks.
β’ Contributions to upstream projects such as Calico, Cilium, Tailscale, or Kubernetes networking SIGs.
β’ Experience managing networking across multiple public clouds and on-premises environments concurrently.
β’ Knowledge of Slurm or other HPC schedulers operating within Kubernetes.
β’ Equity
β’ Comprehensive benefits package
LMI
Sourcegraph
Verra Mobility
Cint
Get handpicked remote jobs straight to your inbox weekly.