Remotery

Network DevOps Engineer, RDMA Fabric Automation

atVultrRemoteUS flagUnited StatesFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$90k – $130k/year

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Automate the deployment and management of extensive RDMA (RoCEv2) Ethernet fabrics across Vultr data centers.

• Create frameworks based on Ansible and Python for provisioning, validating, and resolving issues in underlay and overlay networks.

• Integrate network automation with Vultr’s source-of-truth systems, including NetBox and OpsMill, to enable intent-driven configuration and validation.

• Develop pipelines for telemetry ingestion and correlation utilizing gNMI, Prometheus, Kafka, and custom collectors.

• Collaborate with platform, orchestration, and product engineering teams to enhance RDMA performance, PFC/ECN behavior, and path symmetry throughout fabrics.

• Implement CI/CD workflows for changes in network configurations, including validation, pre-checks, and rollbacks.

• Investigate intricate network behaviors involving flow hashing, congestion domains, ECMP, and overlay interactions.

• Contribute to the design and integration of next-generation GPU and AI interconnect fabrics into Vultr’s global network architecture.


⛳️ Requirements

• Strong comprehension of contemporary data center networking: EVPN-VXLAN, BGP, MLAG, QoS, and traffic engineering.

• Extensive knowledge of RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design.

• Proven experience with automation frameworks such as Ansible.

• Proficiency in Python, Golang, Rust, or PHP.

• Familiarity with telemetry and monitoring stacks like Prometheus, Grafana, Loki, and ELK.

• Prior experience integrating with NetBox, Nautobot, OpsMill, or comparable topology and configuration source-of-truth systems.

• Knowledge of CI/CD systems such as GitHub Actions, Jenkins, and ArgoCD.

• Strong foundation in Linux networking, including namespaces, netlink, and system-level debugging.

• Legally authorized to work in the United States.


🏝️ Benefits

• 100% company-covered insurance premiums for employee medical, dental, and vision plans.

• 401(k) plan with a 100% match up to 4%, featuring immediate vesting.

• Professional Development Reimbursement of $2,500 annually.

• 11 Holidays + Paid Time Off Accrual + Rollover Plan.

• Increased PTO at the 3-year and 10-year anniversaries.

• 1 month of paid sabbatical every 5 years.

• Annual Anniversary Bonus.

• $500 stipend for remote office setup in the first year + $400 in each subsequent year.

• Internet reimbursement up to $75 per month.

• Gym membership reimbursement up to $50 per month.

• Company-paid Wellable subscription.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers