
AI Systems, Model Optimization
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in California.
β’ Create the pathway from model architecture to physical silicon.
β’ Design the training methodologies, optimization techniques, and infrastructure necessary for efficient AI model execution on innovative compute platforms.
β’ Construct detailed performance models to assess trade-offs in compute, memory, and energy usage.
β’ Lead the partitioning and mapping of intricate AI models onto hardware.
β’ Develop and implement Quantization-Aware Training (QAT), noise-aware training, and sparsification methods.
β’ Enhance and optimize kernels utilizing low-level programming frameworks such as CUDA, Triton, or CUTLASS.
β’ Serve as a bridge between AI model architects and hardware/infrastructure engineering teams.
β’ An MS/PhD or equivalent research/project experience in a quantitative discipline, including AI/Machine Learning, Computer Science, Physics, Electrical Engineering, or Applied Mathematics.
β’ Profound practical knowledge of the contemporary AI/ML stack along with optimized compilation and execution of algorithms on current GPU systems.
β’ Demonstrated experience in profiling, identifying, and addressing performance bottlenecks in complex ML codebases.
β’ Proven ability to correlate cutting-edge AI model architectures (such as Transformers, Mixture of Experts, and diffusion models) with system performance outcomes.
β’ Extensive experience with PyTorch, including its internal workings, torch.compile, and distributed data parallel (DDP) / fully sharded data parallel (FSDP) libraries.
β’ A comprehensive package including best-in-class health benefits
β’ 401k matching
β’ Truly unlimited PTO
β’ Complimentary meals in our Palo Alto office
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.