
Staff Network Engineer – AI Fabric, Datacenter and Edge Networking
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
• Design, implement, and manage the network infrastructure that supports a GPU cloud platform for cloud gaming, artificial intelligence, and machine learning applications within telecom carrier networks.
• Create and maintain high-performance GPU networking fabrics and extensive RoCE fabrics.
• Enhance network performance for GPU communication patterns and east-west traffic.
• Establish reusable AI-fabric reference architectures and design principles.
• Develop and manage Layer-2 and Layer-3 datacenter networks utilizing scalable routing architectures.
• Ensure tenant isolation, manage ingress/egress routing, traffic management, and overlay networking.
• Deploy and oversee north-south security infrastructure, WAF protections, and platform security controls.
• Design private interconnects, dark fiber rings, high-capacity WAN connectivity, and global backbone integration.
• Lead end-to-end engineering delivery, from design and lab validation to production deployment.
• Validate network bills of materials, contribute to datacenter layouts and rack elevations, and manage capacity planning.
• Set deployment standards, validation criteria, rollback strategies, and reusable acceptance patterns.
• Take ownership of networking operational performance and reliability.
• Automate provisioning, configuration management, monitoring, and lifecycle management processes.
• Direct incident response and root-cause analysis for significant network events.
• Define and monitor SLAs, SLOs, reliability metrics, and operational benchmarks.
• Collaborate with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams.
• Serve as the primary authority on networking design and influence the platform's networking roadmap.
• Mentor engineers and contribute to the development of the networking function.
• Extensive hands-on experience in designing and managing large-scale datacenter networks.
• Expert-level knowledge of BGP, OSPF, ECMP, and EVPN/VXLAN.
• Proven experience in operating high-speed Ethernet networks within production environments.
• Familiarity with NVIDIA/Mellanox networking platforms.
• Profound expertise in designing and managing networking fabrics for extensive GPU clusters and distributed AI workloads.
• In-depth understanding of NCCL communication patterns.
• Experience in tuning RoCE fabrics.
• Strong knowledge of RDMA transport behavior and potential failure modes.
• Practical experience in implementing and tuning PFC and ECN.
• Understanding of GPU collective communication patterns.
• Experience in designing rail-optimized GPU networking fabrics.
• Awareness of networking performance impacts on PyTorch and TensorFlow.
• Capability to debug cross-layer issues involving hardware, firmware, kernel networking, and distributed application communication layers.
• Strong knowledge of networking hardware, optics, and high-speed interconnects.
• Experience in designing network observability systems.
• Proficient automation skills using Python and/or Bash.
• Experience in applying software engineering practices to infrastructure automation.
• Proven ability to lead complex technical initiatives across different teams.
• Demonstrated capability to set architectural direction and promote the adoption of engineering standards.
• Strong mentoring skills.
• Experience in owning architecture and direct implementation in lean or rapidly scaling environments is highly preferred.
• Competitive compensation package that reflects your skills and experience.
• Flexible working environment.
• Hybrid-friendly work arrangements.
• International and diverse workplace culture.
• Opportunities for career advancement in a rapidly growing scale-up.
Ignít - A Claranet Portugal Company
JetBrains
MedAdvisor Solutions US
LottieFiles
Get handpicked remote jobs straight to your inbox weekly.