
Member of Technical Staff – Datacenter Networking
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in California.
• Design and implement scalable datacenter network architectures for GPU training, inference, storage, and management traffic.
• Configure and manage high-performance Ethernet/RoCE and InfiniBand networks, adhering to standards for routing, redundancy, and capacity.
• Automate processes for network provisioning, configuration validation, upgrades, and rollback.
• Troubleshoot packet loss, congestion, link failures, and collective communication performance across hosts and switches.
• Benchmark overall network performance with infrastructure and ML teams, converting workload requirements into quantifiable acceptance criteria.
• Develop monitoring systems for port health, errors, utilization, congestion, and fabric topology.
• Enhance incident response procedures and runbooks.
• Collaborate with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.
• Engage directly with customers at the forefront of AI innovation.
• Work alongside the engineering team to support systems that drive AI advancements.
• A minimum of 3 years of experience in production datacenter networking.
• Comprehensive knowledge of Ethernet, TCP/IP, routing, switching, and redundant network design.
• Practical experience in high-performance GPU networking utilizing InfiniBand or RoCE.
• Proven ability to troubleshoot networking issues across Linux hosts, NICs, switches, and physical connections.
• Proficiency in automating network operations using Python, Ansible, or similar tools.
• Familiarity with leaf-spine architectures, BGP, ECMP, VLANs, and network segmentation.
• Understanding of RDMA concepts, performance tuning, congestion control, and lossless Ethernet considerations.
• Knowledge of Linux networking, including NIC drivers and firmware, packet capture, and throughput/latency testing.
• Expertise in optics, transceivers, cable management, and link-level diagnostics.
• Awareness of safe change management, configuration versioning, telemetry, and alerting processes.
• Experience managing 400G/800G networks or large multi-rack GPU clusters.
• Familiarity with NVIDIA Spectrum or Quantum networking.
• Experience with NCCL performance analysis and troubleshooting distributed training.
• Knowledge of EVPN/VXLAN, SONiC, or network source-of-truth systems.
• Background in network simulation, automated validation, and capacity planning.
• Equity incentives.
• Remote work option.
PerfectServe
Prime Intellect
Juniper Square
Get handpicked remote jobs straight to your inbox weekly.