
AI Systems, Training
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in California.
• Construct and sustain highly optimized, model-specific training stacks finely tuned for leading-edge generative vision, language, and world models.
• Design and expand multi-node distributed training systems, implementing elastic sharding and resilient data streaming pipelines for rapid, large-scale iteration. Establish robust model checkpointing and recovery processes.
• Develop and refine kernels utilizing low-level programming models such as CUDA and Triton. Create comprehensive benchmarking suites to monitor Model Flops Utilization (MFU), memory bandwidth, and convergence stability.
• Serve as a liaison, engaging in discussions about algorithmic trade-offs with theorists and translating model requirements into detailed specifications for infrastructure and hardware engineering teams.
• An MS/PhD or comparable research/project experience in a quantitative discipline such as AI/Machine Learning, Computer Science, Physics, Electrical Engineering, or Applied Mathematics.
• Experienced with the contemporary ML software stack. Proven capability to correlate state-of-the-art AI model architectures (e.g., transformers, Mixture of Experts, diffusion models) with system performance implications. Extensive expertise in how models are distributed across a cluster, along with mastery of communication primitives and parallelism techniques.
• Demonstrated history of implementing, debugging, and maintaining production-quality training frameworks—such as Megatron-LM, DeepSpeed, Ray, PyTorch Lightning—transforming raw compute into a dependable model-building factory.
• Best-in-class health benefits
• 401k matching
• Truly unlimited PTO
• Complimentary meals in our Palo Alto office
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.