
Engineering Manager, Deep Learning Inference
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in California, +2 more states.
• Lead, mentor, and expand a high-performing engineering team dedicated to deep learning inference and GPU-accelerated software.
• Shape the strategy, roadmap, and implementation of NVIDIA's OSS inference frameworks engineering.
• Collaborate with internal compiler, libraries, and research teams to provide end-to-end optimized inference pipelines across NVIDIA accelerators.
• Supervise performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
• Guide engineers in embracing best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
• Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s overarching AI and software strategies.
• Cultivate a culture of technical excellence, open collaboration, and ongoing innovation.
• MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related discipline.
• Over 6 years of overall software development experience, including more than 3 years in technical leadership or engineering management.
• Strong expertise in C/C++ software design and development; proficiency in Python is an advantage.
• Practical experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
• Demonstrated success in deploying or optimizing deep learning models in production settings.
• Experience leading teams using Agile or collaborative software development methodologies.
• Highly competitive salaries.
• Comprehensive benefits package.
Remote People
Bestow
Virta Health
Space Inch
Get handpicked remote jobs straight to your inbox weekly.