
GPU DC East-West Network SRE Expert, SME
Posted 21 hours ago

Posted 21 hours ago
This is a fully remote position, open to applicants in California, +1 more state.
• Operate InfiniBand and RoCEv2 networks transmitting NCCL traffic across Bitdeer's US data centers.
• Oversee fat-tree, rail-optimized, and dragonfly network architectures for GPU clusters ranging from 100 to 10,000 GPUs.
• Monitor and optimize InfiniBand and RoCE performance, focusing on adaptive routing, congestion management, and traffic isolation.
• Enhance NCCL communication through topology detection, selection of ring/tree algorithms, and configuration of GDR.
• Manage the firmware lifecycle for InfiniBand switches and Host Channel Adapters (HCAs).
• Diagnose issues such as link flaps, symbol errors, packet drops, routing anomalies, and credit stalls.
• Collaborate with Nvidia/Mellanox support on escalations, bug reports, and Return Merchandise Authorizations (RMAs).
• Integrate IB/RoCE telemetry into the AIOps platform’s data collection pipeline.
• Work alongside the platform team to establish link and straggler predictors.
• Transform incidents into labeled datasets for the fault-prediction engine.
• Develop automated remediation workflows from routine mitigations.
• A minimum of 5 years in data center networking, with at least 3 years dedicated to InfiniBand or RoCE fabrics.
• Practical experience in deploying and managing Nvidia/Mellanox InfiniBand switches at scale.
• In-depth knowledge of InfiniBand subnet management, partitioning, and Quality of Service (QoS).
• Experience in deploying RoCEv2, including configuring Priority Flow Control (PFC), Explicit Congestion Notification (ECN), and Data Center Quantized Congestion Notification (DCQCN).
• Proficient in using UFM or similar InfiniBand fabric management tools.
• Familiarity with 400G/800G optics, cabling standards, and best practices in structured cabling.
• Experience in diagnosing InfiniBand/RoCE issues with tools such as ibdiagnet, perfquery, ibstat, and others.
• Understanding of NCCL and its mapping of GPU communication to network topology.
• Experience in telemetry-driven operations, including building dashboards or alerts on RDMA counters at scale, or the capability to define features for a fabric-health model.
• A mindset geared towards runbook-as-code practices.
• Comprehensive health insurance.
• Generous vacation and paid time off policy.
• Opportunities for professional development and training.
• Flexible working hours and remote work options.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.