
Software Engineer, DGX Cloud AI Infrastructure
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in California, +3 more states.
• Initialize, validate, and troubleshoot extensive AI clusters, infrastructure, and end-to-end workloads.
• Set up, optimize, and evaluate AI pre-training, post-training, and inference workloads utilizing PyTorch, NeMo / Megatron, TensorRT-LLM, and related NVIDIA AI software frameworks.
• Conduct root-cause analysis for failures in large distributed environments.
• Contribute to the development of resilience and failure-attribution tools that identify, triage, and assign responsibility for node, fabric, and workload failures throughout the cluster.
• Create and sustain repeatable benchmark suites, automation processes, acceptance criteria, and qualification workflows on new platforms.
• Optimize runtime settings, communication parameters, and deployment configurations in collaboration with framework, systems, and platform teams.
• Provide actionable, data-driven insights based on profiling, benchmark outcomes, and cluster characterization.
• Bachelor’s or Master’s degree in Computer Science or a related technical discipline (or equivalent experience).
• Proven experience in developing software for AI, HPC, or systems-level applications.
• Practical experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
• A background in debugging and scaling distributed systems.
• Experience in troubleshooting and triaging AI applications across the entire stack, from application level to hardware.
• Familiarity with operating workloads in scheduled, containerized cluster environments.
• Exceptional analytical, debugging, and communication skills, with a collaborative approach across teams.
• Proficient in Python and C/C++ programming.
• Direct experience with NCCL and CUDA-aware distributed execution.
• Extensive knowledge of the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging.
• Experience in developing acceptance tests, benchmark harnesses, regression gates, or cluster qualification tools for AI platforms, including MLPerf.
• Familiarity with diagnosing performance jitter.
• Experience in building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.
• Equity
• Benefits
Malbek
ElevenLabs
ElevenLabs
Solidus Labs
Get handpicked remote jobs straight to your inbox weekly.