
Principal HPC Network Engineer – Remote in the EU
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
• Create, implement, and sustain high-performance network infrastructures tailored for HPC environments, emphasizing InfiniBand fabrics.
• Resolve intricate network issues within InfiniBand and Ethernet settings, ensuring minimal downtime and peak performance.
• Oversee and enhance InfiniBand components, such as switches, HCAs, subnet managers, and fabric configurations.
• Conduct performance tuning, monitoring, and capacity planning for HPC networking systems.
• Establish and uphold network security utilizing Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
• Identify and remedy issues related to routing, switching, latency, and throughput within hybrid network environments.
• Collaborate with compute, storage, and platform teams to facilitate HPC workloads and cluster operations.
• Create and maintain documentation for network architecture, configurations, and operational procedures.
• Engage in on-call rotations and offer escalation support for critical incidents.
• Lead or assist in network upgrades, migrations, and new deployments.
• A minimum of 5 years of experience in network engineering, particularly in HPC or data center settings.
• Extensive hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
• Strong grasp of networking principles: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
• Demonstrated experience in deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
• Familiarity with network performance analysis and troubleshooting tools.
• Knowledge of Linux systems and scripting for automation (e.g., Bash, Python).
• Excellent analytical and problem-solving capabilities.
• Preferred: Experience with large-scale HPC clusters or AI/ML infrastructure.
• Preferred: Understanding of RDMA, MPI, and concepts related to low-latency networking.
• Preferred: Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
• Preferred: Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).
• Operate within some of the most advanced AI infrastructure environments currently in production.
• Work with cutting-edge NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to defining operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the adoption of AI-powered operational capabilities via k0rdent AI.
• Collaborate with highly skilled engineers to tackle complex infrastructure and platform challenges at scale.
• Join a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
Mirantis
blueAPACHE
CVS Health
Humana
Get handpicked remote jobs straight to your inbox weekly.