
Software Engineer, DGX Cloud AI Infrastructure
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California, +3 more states.
• Set up, validate, and troubleshoot large-scale AI clusters, infrastructure, and end-to-end workloads.
• Establish, optimize, and evaluate AI pre-training, post-training, and inference tasks utilizing PyTorch, NeMo / Megatron, TensorRT-LLM, and related NVIDIA AI software ecosystems.
• Conduct root-cause analysis for failures in extensive distributed environments.
• Contribute to resilience and failure-attribution tools that identify, manage, and assign node, fabric, and workload failures throughout the cluster.
• Develop and sustain repeatable benchmarking suites, automation processes, acceptance criteria, and qualification workflows on new platforms.
• Optimize runtime settings, communication parameters, and deployment configurations in collaboration with framework, systems, and platform teams.
• Provide actionable, data-driven insights based on profiling, benchmarking results, and cluster characterization.
• Bachelor’s or Master’s degree in Computer Science or a related technical field, or equivalent experience.
• Over 3 years of experience in developing software for AI, HPC, or systems-level applications.
• Practical experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
• Background in debugging and scaling distributed systems.
• Proficient in troubleshooting and managing AI applications across the entire stack, from application level to hardware.
• Experience in operating workloads in scheduled, containerized cluster environments.
• Exceptional analytical, debugging, and communication skills, with a team-oriented approach.
• Strong programming skills in Python and C/C++.
• Practical experience with NCCL and CUDA-aware distributed execution.
• Knowledge of the RDMA software stack, including NCCL, IB verbs, UCX, and libfabric.
• Understanding of InfiniBand / RoCE congestion debugging.
• Experience in creating acceptance tests, benchmark harnesses, regression gates, or cluster qualification tools for AI platforms, including MLPerf.
• Familiarity with diagnosing performance jitter.
• Experience in developing resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.
• Equity
• Benefits
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.