Senior Software Engineer – Cluster Networking

Posted 2 days ago

This is a fully remote position, open to applicants in California, +4 more states.

πŸ“‹ Description

β€’ Take ownership of and advance the Kubernetes networking framework for GPU clusters operating at a scale of several thousand nodes.

β€’ Architect, manage, and enhance the overlay network, including CNI, mesh, and VPN topologies (Tailscale, WireGuard), as well as the gateways that link control and data planes.

β€’ Design, manage, and enhance L7 gateways/load balancers/tunnels (Envoy, Cloudflare).

β€’ Identify and resolve scalability limitations such as packet loss under heavy load, control-plane saturation, IP address management exhaustion, and failure modes when exceeding a few thousand nodes.

β€’ Develop scale-testing environments and validation frameworks to detect networking regressions prior to production deployment.

β€’ Troubleshoot complex and ambiguous issues throughout the stack, tracing symptoms in Slurm or training jobs back to their networking root causes.

β€’ Collaborate with cloud and neocloud providers to define network topology, requirements, and capabilities for new clusters.

β€’ Offer senior-level technical expertise to a distributed team within the Cluster Networking domain.


⛳️ Requirements

β€’ BS/MS in Computer Science, Electrical Engineering, or a related discipline, or equivalent professional experience.

β€’ Over 6 years of professional experience in systems, networking, or infrastructure software engineering.

β€’ Extensive knowledge of Kubernetes networking architecture and CNI standards, with hands-on experience operating Calico highly preferred.

β€’ Expertise in designing and maintaining contemporary mesh and VPN networking topologies such as Tailscale, WireGuard, or similar solutions.

β€’ Strong foundational knowledge of Linux networking: routing, netfilter and iptables/nftables, packet marking, network namespaces, and their interaction with container runtimes.

β€’ Proven ability to debug distributed network issues at scale, including packet capture, tracing, and correlating behavior across numerous hosts to identify a single root cause.

β€’ Proficiency in Go, Python, C, or a similar systems programming language.

β€’ Excellent written and verbal communication skills, with the capacity to collaborate effectively with engineers across various time zones.

β€’ Direct experience in architecting and managing large-scale Kubernetes configurations involving thousands of concurrent nodes.

β€’ Familiarity with high-performance fabrics in AI or HPC settings, including InfiniBand, RoCE, or RDMA over converged networks.

β€’ Contributions to upstream projects such as Calico, Cilium, Tailscale, or Kubernetes networking SIGs.

β€’ Experience managing networking across multiple public clouds and on-premises environments concurrently.

β€’ Knowledge of Slurm or other HPC schedulers operating within Kubernetes.


🏝️ Benefits

β€’ Equity

β€’ Comprehensive benefits package

People also viewed

LMI11 hours ago

Full Stack Developer

US flagUnited States OnlyFull-timeFull-stack Engineer$122.2k – $211.3k/year
ApplyView job
Sourcegraph11 hours ago

Tech Lead – Code Plane

EuropeFull-timeFull-stack Engineer$144k – $192k/year
ApplyView job
Verra Mobility12 hours ago

Vice President, Software Engineering

US flagTexas OnlyFull-timeFull-stack Engineer
ApplyView job
Cint12 hours ago

Staff Software Engineer – DSM Team

ES flagSpain OnlyFull-timeFull-stack Engineer
ApplyView job
CI&T12 hours ago

Mid Level FullStack Developer – .NET, React

BR flagBrazil OnlyFull-timeFull-stack EngineerR$1 – R$2/month
ApplyView job
Akamai Technologies12 hours ago

Senior Software Engineer

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers