
Principal HPC Network Engineer
Posted Jul 7

Posted Jul 7
This is a fully remote position, open to applicants in Europe.
• Design, implement, and sustain high-performance network infrastructures tailored for HPC environments, emphasizing InfiniBand fabrics.
• Resolve intricate network issues across both InfiniBand and Ethernet settings, ensuring minimal downtime and peak performance.
• Oversee and enhance InfiniBand components, which include switches, HCAs, subnet managers, and fabric configurations.
• Conduct performance tuning, monitoring, and capacity planning for HPC networking systems.
• Establish and uphold network security utilizing Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
• Identify and rectify issues related to routing, switching, latency, and throughput within hybrid network environments.
• Collaborate with compute, storage, and platform teams to facilitate HPC workloads and cluster operations.
• Create and maintain documentation detailing network architecture, configurations, and operational procedures.
• Engage in on-call rotations and offer escalation support for critical incidents.
• Lead or assist in network upgrades, migrations, and new deployments.
• Over 5 years of experience in network engineering, specifically in HPC or data center environments.
• Extensive hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
• Strong grasp of networking fundamentals, including TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
• Proven expertise in deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
• Experience with network performance analysis and troubleshooting tools.
• Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
• Excellent analytical and problem-solving abilities.
• Preferred experience with large-scale HPC clusters or AI/ML infrastructures.
• Preferred knowledge of RDMA, MPI, and low-latency networking concepts.
• Preferred certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
• Preferred experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).
• Strong communication and collaboration skills.
• Ability to work independently and tackle complex technical challenges.
• Detail-oriented with a proactive approach to resolving problems.
• Operate within some of the most advanced AI infrastructure environments currently in production.
• Work with cutting-edge NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to defining operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the integration of AI-powered operational capabilities through k0rdent AI.
• Collaborate with highly skilled engineers to address complex infrastructure and platform challenges at scale.
• Join a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
Koniag Government Services
BPCS, Comprehensive marketing solutions, ltd.
Titan Technologies
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.