
Principal HPC Network Engineer
Posted Jul 8

Posted Jul 8
This is a fully remote position, open to applicants in Europe.
• Design, implement, and sustain high-performance networking infrastructures tailored for HPC environments, emphasizing InfiniBand fabrics.
• Diagnose intricate network problems within both InfiniBand and Ethernet settings, guaranteeing minimal downtime and peak performance.
• Oversee and enhance InfiniBand components, which include switches, HCAs, subnet managers, and fabric configurations.
• Conduct performance tuning, monitoring, and capacity planning for HPC networking systems.
• Apply and uphold network security utilizing Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
• Identify and rectify issues related to routing, switching, latency, and throughput across hybrid network landscapes.
• Collaborate with compute, storage, and platform teams to facilitate HPC workloads and cluster operations.
• Create and maintain documentation for network architecture, configurations, and operational procedures.
• Engage in on-call rotations and offer escalation support for critical incidents.
• Lead or assist in network upgrades, migrations, and new implementations.
• Over 5 years of experience in network engineering, particularly in HPC or data center environments.
• Extensive hands-on expertise with InfiniBand technologies (e.g., Mellanox/NVIDIA).
• Strong knowledge of networking principles: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network architecture.
• Demonstrated experience in deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
• Proficiency in network performance analysis and troubleshooting tools.
• Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
• Exceptional analytical and problem-solving abilities.
• Experience with large-scale HPC clusters or AI/ML infrastructure is preferred.
• Knowledge of RDMA, MPI, and concepts of low-latency networking is preferred.
• Relevant certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent are preferred.
• Operate within some of the most cutting-edge AI infrastructure environments currently in production.
• Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to establishing operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the integration of AI-powered operational capabilities through k0rdent AI.
• Collaborate with highly skilled engineers tackling complex infrastructure and platform challenges at scale.
• Join a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
Thinkahead Consultant Psychologist Pty Ltd
Zayo Group
Thinkahead Consultant Psychologist Pty Ltd
Mirantis
Get handpicked remote jobs straight to your inbox weekly.