Remotery

Senior Principal AI Engineer

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Design and manage distributed training systems for extensive neural networks across GPU clusters.

• Enhance multi-node, multi-GPU performance for optimal throughput and resource utilization.

• Identify and resolve issues related to compute, memory, and networking constraints.

• Advance training stability and fault tolerance at scale.

• Collaborate with research and applied ML teams to implement large-model training pipelines into production.

• Construct and refine GPU cluster orchestration utilizing Slurm, Kubernetes, Ray, and RunAI.

• Guarantee effective scheduling, isolation, and fairness among training tasks.

• Fine-tune and troubleshoot distributed communication leveraging NCCL, RDMA, InfiniBand, and NVLink.

• Expand large-model training capabilities using PyTorch Distributed, Megatron-LM, and DeepSpeed.

• Oversee multi-node launch configurations, recovery from failures, and performance enhancements.

• Implement activation checkpointing, ZeRO Stage 1–3, and various offloading strategies.

• Balance compute, memory, and communication to support increased model size and batch scaling.

• Address issues related to GPU usage, networking, communication, instability, and large-scale training failures.

• Facilitate larger models, expedited iteration cycles, and more dependable research-to-production pipelines.


⛳️ Requirements

• Extensive hands-on experience with distributed systems or machine learning systems.

• Proven experience executing large-scale workloads on GPU clusters.

• Practical experience with PyTorch distributed training in a production environment.

• Strong grasp of data, tensor, and pipeline parallelism concepts.

• Fundamental understanding of GPU communication and networking protocols.

• Familiarity with Slurm, Kubernetes, Ray, and RunAI platforms.

• Knowledge of NCCL, RDMA, InfiniBand, and NVLink technologies.

• Experience with PyTorch Distributed, Megatron-LM, and DeepSpeed frameworks.

• Proficiency in activation checkpointing, ZeRO Stage 1–3, and offloading methodologies.

• Experience with large language models or foundational models is highly desirable.

• Strong systems expertise is prioritized over mere model architecture experience.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health and wellness benefits.

• Opportunities for professional development and continuous learning.

• Flexible work arrangements and a supportive work environment.

People also viewed

dexter health9 hours ago

Applied AI Engineer

DE flagGermany OnlyFull-timeAI Engineer
ApplyView job
Blend3609 hours ago

Lead AI Engineer – Agentic Engineering

IN flagIndia OnlyFull-timeAI Engineer
ApplyView job
CI&T9 hours ago

Senior Generative AI Developer

BR flagBrazil OnlyFull-timeAI Engineer
ApplyView job
Lingaro10 hours ago

ML/AI Engineer

PL flagPoland OnlyFreelanceAI Engineer
ApplyView job
Orion Innovation10 hours ago

Cloud/AI Developer

US flagUnited States OnlyFull-timeAI Engineer
ApplyView job
Tango11 hours ago

Senior Applied AI Engineer

US flagUnited States, +1 more stateFull-timeAI Engineer$160k – $190k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers