
Senior Deep Learning Algorithm Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in California.
• Design, develop, and enhance workloads for NVIDIA’s Megatron Core and NeMo Framework teams.
• Broaden the capabilities of Megatron Core and NeMo Framework for the development, training, and optimization of LLM and multimodal foundation models.
• Create and implement distributed training algorithms, model parallel paradigms, and optimizations for models.
• Establish robust APIs while analyzing and tuning performance metrics.
• Enrich toolkits and libraries to ensure they are more comprehensive and cohesive.
• Collaborate with internal partners, users, and the open-source community to design and implement optimized solutions.
• Develop algorithms applicable to AI/deep learning, data analytics, machine learning, or scientific computing.
• Contribute to and enhance the open-source NeMo-RL, Megatron Core, and NeMo Framework.
• Address large-scale, end-to-end AI training and inference challenges, including orchestration, data preprocessing, training, tuning, and deployment.
• Enhance model architectures, distributed training algorithms, and model parallel paradigms.
• Tune performance and optimize model training and fine-tuning using mixed-precision recipes on next-generation NVIDIA GPU architectures.
• Research, prototype, and develop robust, scalable AI tools and pipelines.
• MS, PhD, or equivalent experience in Computer Science, AI, Applied Mathematics, or related fields.
• Over 5 years of industry experience.
• Familiarity with AI frameworks such as PyTorch, JAX, or Ray, as well as inference and deployment environments like TRTLLM, vLLM, or SGLang.
• Proficient in Python programming, software design, debugging, performance analysis, test design, and documentation.
• Proven track record of effectively collaborating on multiple engineering initiatives and enhancing AI libraries with innovative solutions.
• Strong grasp of AI/deep learning fundamentals along with their practical applications.
• Practical experience in large-scale AI training and a solid understanding of compute system concepts, including latency/throughput bottlenecks, pipelining, and multiprocessing.
• Experience with reinforcement learning algorithms and compute patterns.
• Knowledge in distributed computing, model parallelism, and mixed-precision training.
• Familiarity with generative AI techniques as they relate to LLM and multimodal learning involving text, image, and video.
• Understanding of GPU/CPU architecture and associated numerical software.
• Equity
• Benefits
Anduril Industries
Sargent & Lundy
Sargent & Lundy
Get handpicked remote jobs straight to your inbox weekly.