Remotery

Forward Deployed Engineer – SRE

Posted Aug 1

This is a fully remote position, open to applicants in California.

πŸ“‹ Description

β€’ Act as the main technical liaison for teams managing extensive training and inference tasks.

β€’ Collaborate within customer settings to troubleshoot actual failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, and OOM patterns.

β€’ Analyze and enhance distributed training performance on active workloads.

β€’ Take ownership of reliability outcomes for the accounts assigned to you.

β€’ Monitor the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink).

β€’ Lead incident response efforts for intricate, multi-layer failures involving hardware, networking, orchestration, and ML frameworks.

β€’ Transform every recurring deployment issue into automated solutions.


⛳️ Requirements

β€’ Practical experience in operating GPU clusters in production environments (NVIDIA A100/H100/H200/B200 or equivalent).

β€’ Hands-on experience with InfiniBand, RoCE, or NVLink fabrics related to distributed training.

β€’ Understanding of how large training and inference jobs function in practice.

β€’ Advanced Linux skills: including kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at both the syscall and hardware levels.

β€’ Significant experience running Kubernetes in production with GPU workloads.

β€’ Strong engineering expertise in Python, Go, or Bash.

β€’ Proficiency in Infrastructure-as-Code (Terraform, Helm, Ansible, or equivalent).

β€’ Experience in developing monitoring and alerting systems for GPU-specific telemetry.


🏝️ Benefits

β€’ Comprehensive benefits package: covering you and your dependents, including healthcare, dental, and vision insurance.

β€’ 401(k) plan.

β€’ Unlimited paid time off (PTO).

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers