Remotery

Software Engineer, DGX Cloud AI Infrastructure

Posted 2 days ago

This is a fully remote position, open to applicants in California, +3 more states.

📋 Description

• Set up, validate, and troubleshoot large-scale AI clusters, infrastructure, and end-to-end workloads.

• Establish, optimize, and evaluate AI pre-training, post-training, and inference tasks utilizing PyTorch, NeMo / Megatron, TensorRT-LLM, and related NVIDIA AI software ecosystems.

• Conduct root-cause analysis for failures in extensive distributed environments.

• Contribute to resilience and failure-attribution tools that identify, manage, and assign node, fabric, and workload failures throughout the cluster.

• Develop and sustain repeatable benchmarking suites, automation processes, acceptance criteria, and qualification workflows on new platforms.

• Optimize runtime settings, communication parameters, and deployment configurations in collaboration with framework, systems, and platform teams.

• Provide actionable, data-driven insights based on profiling, benchmarking results, and cluster characterization.


⛳️ Requirements

• Bachelor’s or Master’s degree in Computer Science or a related technical field, or equivalent experience.

• Over 3 years of experience in developing software for AI, HPC, or systems-level applications.

• Practical experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.

• Background in debugging and scaling distributed systems.

• Proficient in troubleshooting and managing AI applications across the entire stack, from application level to hardware.

• Experience in operating workloads in scheduled, containerized cluster environments.

• Exceptional analytical, debugging, and communication skills, with a team-oriented approach.

• Strong programming skills in Python and C/C++.

• Practical experience with NCCL and CUDA-aware distributed execution.

• Knowledge of the RDMA software stack, including NCCL, IB verbs, UCX, and libfabric.

• Understanding of InfiniBand / RoCE congestion debugging.

• Experience in creating acceptance tests, benchmark harnesses, regression gates, or cluster qualification tools for AI platforms, including MLPerf.

• Familiarity with diagnosing performance jitter.

• Experience in developing resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.


🏝️ Benefits

• Equity

• Benefits

People also viewed

Cloudera7 hours ago

Staff Software Engineer, Flink/Streaming Analytics

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job
Stellar Cyber7 hours ago

Senior / Staff Software Engineer – Parser Team

US flagUnited States OnlyFull-timeFull-stack Engineer$150k – $200k/year
ApplyView job
Pragmatike8 hours ago

Senior Founding Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike8 hours ago

Staff Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike8 hours ago

Founding Product Engineer – YCombinator Experience

US flagCalifornia OnlyFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike8 hours ago

Lead Product Engineer

US flagCalifornia, +2 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers