
Forward Deployed Engineer β SRE
Posted Aug 1

Posted Aug 1
This is a fully remote position, open to applicants in California.
β’ Act as the main technical liaison for teams managing extensive training and inference tasks.
β’ Collaborate within customer settings to troubleshoot actual failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, and OOM patterns.
β’ Analyze and enhance distributed training performance on active workloads.
β’ Take ownership of reliability outcomes for the accounts assigned to you.
β’ Monitor the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink).
β’ Lead incident response efforts for intricate, multi-layer failures involving hardware, networking, orchestration, and ML frameworks.
β’ Transform every recurring deployment issue into automated solutions.
β’ Practical experience in operating GPU clusters in production environments (NVIDIA A100/H100/H200/B200 or equivalent).
β’ Hands-on experience with InfiniBand, RoCE, or NVLink fabrics related to distributed training.
β’ Understanding of how large training and inference jobs function in practice.
β’ Advanced Linux skills: including kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at both the syscall and hardware levels.
β’ Significant experience running Kubernetes in production with GPU workloads.
β’ Strong engineering expertise in Python, Go, or Bash.
β’ Proficiency in Infrastructure-as-Code (Terraform, Helm, Ansible, or equivalent).
β’ Experience in developing monitoring and alerting systems for GPU-specific telemetry.
β’ Comprehensive benefits package: covering you and your dependents, including healthcare, dental, and vision insurance.
β’ 401(k) plan.
β’ Unlimited paid time off (PTO).
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.