Software Engineer, DGX Cloud AI Infrastructure

Posted 5 days ago

This is a fully remote position, open to applicants in California, +3 more states.

📋 Description

• Initialize, validate, and troubleshoot extensive AI clusters, infrastructure, and end-to-end workloads.

• Set up, optimize, and evaluate AI pre-training, post-training, and inference workloads utilizing PyTorch, NeMo / Megatron, TensorRT-LLM, and related NVIDIA AI software frameworks.

• Conduct root-cause analysis for failures in large distributed environments.

• Contribute to the development of resilience and failure-attribution tools that identify, triage, and assign responsibility for node, fabric, and workload failures throughout the cluster.

• Create and sustain repeatable benchmark suites, automation processes, acceptance criteria, and qualification workflows on new platforms.

• Optimize runtime settings, communication parameters, and deployment configurations in collaboration with framework, systems, and platform teams.

• Provide actionable, data-driven insights based on profiling, benchmark outcomes, and cluster characterization.


⛳️ Requirements

• Bachelor’s or Master’s degree in Computer Science or a related technical discipline (or equivalent experience).

• Proven experience in developing software for AI, HPC, or systems-level applications.

• Practical experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.

• A background in debugging and scaling distributed systems.

• Experience in troubleshooting and triaging AI applications across the entire stack, from application level to hardware.

• Familiarity with operating workloads in scheduled, containerized cluster environments.

• Exceptional analytical, debugging, and communication skills, with a collaborative approach across teams.

• Proficient in Python and C/C++ programming.

• Direct experience with NCCL and CUDA-aware distributed execution.

• Extensive knowledge of the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging.

• Experience in developing acceptance tests, benchmark harnesses, regression gates, or cluster qualification tools for AI platforms, including MLPerf.

• Familiarity with diagnosing performance jitter.

• Experience in building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.


🏝️ Benefits

• Equity

• Benefits

People also viewed

Malbek17 hours ago

Full Stack Developer

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
ElevenLabs19 hours ago

Forward Deployed Engineer – Software Engineer

MX flagMexico OnlyFull-timeFull-stack Engineer
ApplyView job
ElevenLabs21 hours ago

Forward Deployed Engineer – Software Engineer

GB flagUnited Kingdom OnlyFull-timeFull-stack Engineer
ApplyView job
Solidus Labs21 hours ago

Tech Lead, Transaction Monitoring

GB flagUnited Kingdom OnlyFull-timeFull-stack Engineer
ApplyView job
Marigold1 day ago

Full Stack Software Engineer

AU flagAustralia OnlyFull-timeFull-stack Engineer
ApplyView job
NationsBenefits1 day ago

Software Development Engineer II

US flagFlorida OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers