Remotery

GPU DC East-West Network SRE Expert, SME

Posted 21 hours ago

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Operate InfiniBand and RoCEv2 networks transmitting NCCL traffic across Bitdeer's US data centers.

• Oversee fat-tree, rail-optimized, and dragonfly network architectures for GPU clusters ranging from 100 to 10,000 GPUs.

• Monitor and optimize InfiniBand and RoCE performance, focusing on adaptive routing, congestion management, and traffic isolation.

• Enhance NCCL communication through topology detection, selection of ring/tree algorithms, and configuration of GDR.

• Manage the firmware lifecycle for InfiniBand switches and Host Channel Adapters (HCAs).

• Diagnose issues such as link flaps, symbol errors, packet drops, routing anomalies, and credit stalls.

• Collaborate with Nvidia/Mellanox support on escalations, bug reports, and Return Merchandise Authorizations (RMAs).

• Integrate IB/RoCE telemetry into the AIOps platform’s data collection pipeline.

• Work alongside the platform team to establish link and straggler predictors.

• Transform incidents into labeled datasets for the fault-prediction engine.

• Develop automated remediation workflows from routine mitigations.


⛳️ Requirements

• A minimum of 5 years in data center networking, with at least 3 years dedicated to InfiniBand or RoCE fabrics.

• Practical experience in deploying and managing Nvidia/Mellanox InfiniBand switches at scale.

• In-depth knowledge of InfiniBand subnet management, partitioning, and Quality of Service (QoS).

• Experience in deploying RoCEv2, including configuring Priority Flow Control (PFC), Explicit Congestion Notification (ECN), and Data Center Quantized Congestion Notification (DCQCN).

• Proficient in using UFM or similar InfiniBand fabric management tools.

• Familiarity with 400G/800G optics, cabling standards, and best practices in structured cabling.

• Experience in diagnosing InfiniBand/RoCE issues with tools such as ibdiagnet, perfquery, ibstat, and others.

• Understanding of NCCL and its mapping of GPU communication to network topology.

• Experience in telemetry-driven operations, including building dashboards or alerts on RDMA counters at scale, or the capability to define features for a fabric-health model.

• A mindset geared towards runbook-as-code practices.


🏝️ Benefits

• Comprehensive health insurance.

• Generous vacation and paid time off policy.

• Opportunities for professional development and training.

• Flexible working hours and remote work options.

People also viewed

CVS Health6 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam7 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion7 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios7 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart8 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain8 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers