
Senior Software Engineer, CUDA Deep Learning Systems
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in California, +1 more state.
• Investigate, analyze, and prototype system enhancements for cutting-edge deep learning models that bridge high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
• Design and optimize distributed computing architectures, ranging from single-node setups to cluster-scale supercomputing environments.
• Create, implement, and enhance custom high-performance CUDA kernels tailored for new neural network architectures and workloads.
• Examine hardware-software interactions to pinpoint and address performance bottlenecks in training and inference workflows.
• Collaborate with AI researchers, hardware and software architects, kernel and compiler developers, and CUDA driver specialists to co-design systems and algorithms.
• Develop exploratory tools and runtime systems aimed at profiling and accelerating innovative deep learning paradigms.
• Produce clean, efficient, and maintainable code, transitioning prototypes into open-source releases, framework integrations, internal tools, or commercial products.
• BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent experience.
• 8+ years of pertinent industry experience or comparable academic experience following degree completion.
• Strong expertise in C++ and Python programming languages.
• Solid understanding of deep learning principles, particularly focusing on transformers.
• Comprehensive knowledge of distributed computing, multi-node scaling, and performance challenges at the cluster scale.
• Demonstrated experience in systems programming, computer architecture, and low-level systems performance optimization.
• Familiarity with GPU accelerator architectures.
• Practical experience with CUDA programming, kernel optimization, and workload profiling.
• Experience in profiling and optimizing generative AI models, including large language models.
• Research experience in machine learning systems or related fields.
• Experience in profiling and optimizing vision models, generative AI architectures, or diffusion models.
• Proven track record of initiative and eagerness to tackle problems across the technology stack.
• Preferred: expertise in the performance internals and execution graphs of PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron.
• Preferred: experience with NCCL, MPI, UCX, and distributed machine learning methodologies such as pipeline, tensor, or expert parallelism.
• Preferred: knowledge of numerical methods and low-precision arithmetic, including NVFP4, MXFP4, FP8, or INT8.
• Preferred: background in deep learning compilers and ML systems, such as Triton, XLA, or torch.compile.
• Preferred: experience in designing agentic AI systems for complex systems and infrastructure challenges.
• Equity
• Benefits
• Equal opportunity employer
• Inclusive work environment
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.