
Principal HPC Network Engineer, remote in the EU
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
β’ Design, implement, and sustain high-performance network infrastructures tailored for HPC environments, with a particular emphasis on InfiniBand fabrics.
β’ Diagnose intricate network problems across both InfiniBand and Ethernet environments, ensuring minimal downtime and peak performance.
β’ Oversee and enhance InfiniBand components, such as switches, HCAs, subnet managers, and fabric configurations.
β’ Conduct performance tuning, monitoring, and capacity planning for HPC networking systems.
β’ Establish and uphold network security utilizing Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
β’ Identify and resolve challenges related to routing, switching, latency, and throughput in hybrid network setups.
β’ Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
β’ Create and maintain documentation for network architecture, configurations, and operational procedures.
β’ Participate in on-call rotations and offer escalation support for critical incidents.
β’ Lead or contribute to network upgrades, migrations, and new deployments.
β’ Over 5 years of experience in network engineering, specifically in HPC or data center environments.
β’ Extensive hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
β’ Strong grasp of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
β’ Demonstrated experience in deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
β’ Familiarity with network performance analysis and troubleshooting tools.
β’ Knowledge of Linux systems and scripting for automation (e.g., Bash, Python).
β’ Excellent analytical and problem-solving abilities.
β’ Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent are advantageous.
β’ Work within some of the most advanced AI infrastructure environments currently in production.
β’ Engage with cutting-edge NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
β’ Contribute to defining operational standards and reliability practices for next-generation AI infrastructure services.
β’ Influence the integration of AI-powered operational capabilities through k0rdent AI.
β’ Collaborate with highly skilled engineers to address complex infrastructure and platform challenges at scale.
β’ Become part of a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
blueAPACHE
Mirantis
CVS Health
Humana
Get handpicked remote jobs straight to your inbox weekly.