
Senior Deep Learning Engineer, CUDA
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in California, +1 more state.
• Integrate cutting-edge CUDA features and runtime abstractions into AI frameworks, progressing from proof of concept to performance analysis and eventual production.
• Examine AI workloads and frameworks to uncover lower-level requirements and opportunities for innovation.
• Engage directly with teams focusing on the latest AI models.
• Enhance the AI compiler-runtime interface for solutions involving multiple GPUs and nodes.
• Create fault-tolerant and flexible solutions tailored for large-scale or dynamic AI workloads.
• Shape the core CUDA roadmap for next-generation deep learning frameworks.
• Work collaboratively across various time zones with AI researchers, hardware and software architects, kernel and compiler developers, and CUDA driver specialists.
• Co-design systems and frameworks that boost performance and programmability.
• Develop innovative tools and runtime systems to profile and expedite new deep learning paradigms.
• Write clean, efficient, maintainable code, transitioning prototypes into open-source releases, framework integrations, internal tools, or commercial products.
• BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
• Over 8 years of relevant industry experience or comparable academic experience following degree completion.
• Development experience with deep learning frameworks such as PyTorch and JAX, along with inference engines like TRT-LLM, vLLM, and SGLang.
• Proficiency in rapid prototyping and development using Python, C++, CUDA, or similar domain-specific languages.
• Strong comprehension of AI models, parallelism, and/or compiler technologies such as torch.compile.
• Experience in performance benchmarking on AI clusters.
• Familiarity with at least one performance profiling toolchain, like PyTorch Profiler or NVIDIA Nsight Systems.
• Understanding of HPC/AI communication concepts.
• Knowledge of computer system architecture, hardware-software interactions, and operating system principles.
• Flexibility and eagerness to learn new frameworks and tools.
• Capability to work and communicate effectively across diverse teams and time zones.
• Preferred additional expertise in the performance internals of deep learning frameworks and execution graphs.
• Preferred hands-on experience with CUDA, NCCL, MPI, UCX, and distributed machine learning methodologies.
• Preferred knowledge in training, distributed inference, mixture-of-experts, reinforcement learning, or kernel authoring with CUDA, Triton, or cuTe.
• Preferred background in deep learning compilers such as Triton, XLA, and torch.compile.
• Preferred experience in programming compute and communication overlap in distributed runtimes.
• Equity
• Benefits
NVIDIA
SentiLink
SentiLink
Leega
Get handpicked remote jobs straight to your inbox weekly.