
Principal HPC Network Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Europe.
• Design, implement, and sustain high-performance network infrastructures for HPC environments, emphasizing InfiniBand fabrics.
• Diagnose intricate network issues within InfiniBand and Ethernet settings, ensuring minimal downtime and peak performance.
• Oversee and enhance InfiniBand components, such as switches, HCAs, subnet managers, and fabric configurations.
• Conduct performance tuning, monitoring, and capacity planning for HPC networking systems.
• Establish and uphold network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
• Identify and resolve issues related to routing, switching, latency, and throughput in hybrid network environments.
• Collaborate with compute, storage, and platform teams to facilitate HPC workloads and cluster operations.
• Create and maintain documentation for network architecture, configurations, and operational protocols.
• Engage in on-call rotations and offer escalation support for critical incidents.
• Lead or contribute to network enhancements, migrations, and new installations.
• 5+ years of experience in network engineering, particularly within HPC or data center settings.
• Extensive hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
• Strong grasp of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
• Demonstrated experience in deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
• Proficiency with network performance analysis and troubleshooting tools.
• Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
• Strong analytical and problem-solving capabilities.
• Preferred: Experience with large-scale HPC clusters or AI/ML infrastructure.
• Knowledge of RDMA, MPI, and low-latency networking concepts.
• Relevant certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
• Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).
• Excellent communication and collaboration abilities.
• Capability to work independently and tackle complex technical challenges.
• Detail-oriented with a proactive approach to problem-solving.
• Operate within some of the most advanced AI infrastructure environments currently in production.
• Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to the definition of operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the adoption of AI-driven operational capabilities through k0rdent AI.
• Collaborate with highly skilled engineers to solve complex infrastructure and platform challenges at scale.
• Join a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
• Competitive salary and benefits package.
TRAC Recruiting
Circle
Astreya
Vamonos IT LLC
Get handpicked remote jobs straight to your inbox weekly.