
Senior Principal AI Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Design and manage distributed training systems for extensive neural networks across GPU clusters.
• Enhance multi-node, multi-GPU performance for optimal throughput and resource utilization.
• Identify and resolve issues related to compute, memory, and networking constraints.
• Advance training stability and fault tolerance at scale.
• Collaborate with research and applied ML teams to implement large-model training pipelines into production.
• Construct and refine GPU cluster orchestration utilizing Slurm, Kubernetes, Ray, and RunAI.
• Guarantee effective scheduling, isolation, and fairness among training tasks.
• Fine-tune and troubleshoot distributed communication leveraging NCCL, RDMA, InfiniBand, and NVLink.
• Expand large-model training capabilities using PyTorch Distributed, Megatron-LM, and DeepSpeed.
• Oversee multi-node launch configurations, recovery from failures, and performance enhancements.
• Implement activation checkpointing, ZeRO Stage 1–3, and various offloading strategies.
• Balance compute, memory, and communication to support increased model size and batch scaling.
• Address issues related to GPU usage, networking, communication, instability, and large-scale training failures.
• Facilitate larger models, expedited iteration cycles, and more dependable research-to-production pipelines.
• Extensive hands-on experience with distributed systems or machine learning systems.
• Proven experience executing large-scale workloads on GPU clusters.
• Practical experience with PyTorch distributed training in a production environment.
• Strong grasp of data, tensor, and pipeline parallelism concepts.
• Fundamental understanding of GPU communication and networking protocols.
• Familiarity with Slurm, Kubernetes, Ray, and RunAI platforms.
• Knowledge of NCCL, RDMA, InfiniBand, and NVLink technologies.
• Experience with PyTorch Distributed, Megatron-LM, and DeepSpeed frameworks.
• Proficiency in activation checkpointing, ZeRO Stage 1–3, and offloading methodologies.
• Experience with large language models or foundational models is highly desirable.
• Strong systems expertise is prioritized over mere model architecture experience.
• Competitive salary and performance-based bonuses.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements and a supportive work environment.
Blend360
Get handpicked remote jobs straight to your inbox weekly.