
Senior Deep Learning Compiler Engineer – XLA
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in California, +2 more states.
• Create optimization algorithms for compilers tailored for deep learning tasks.
• Enhance both inference and training performance for the JAX framework and OpenXLA compiler on NVIDIA GPUs at scale.
• Collaborate with partners in deep learning frameworks and teams focused on hardware architecture.
• Develop and apply compiler optimization strategies for deep learning network graphs.
• Design techniques for graph partitioning and tensor sharding to facilitate distributed training and inference.
• Conduct performance tuning and thorough analysis.
• Execute code generation for NVIDIA GPU backends utilizing MLIR, LLVM, and OpenAI Triton.
• Create user-facing features within JAX and its associated libraries.
• Engage in general software engineering tasks.
• Partner with GPU hardware engineering teams to devise AI compiler software features for future-generation GPUs.
• Establish project objectives and scope while leading development initiatives.
• Implement software engineering and testing methodologies.
• Guide junior engineers and interns through mentorship.
• A Bachelor's, Master's, or Ph.D. in Computer Science, Computer Engineering, or a related field, or equivalent experience.
• At least 4 years of relevant work or research experience focused on performance analysis and compiler optimizations.
• Ability to operate independently, define project goals and scope, and lead development projects.
• Proficient in clean software engineering and testing practices.
• Exceptional skills in C/C++ programming and software design.
• Experience in debugging, performance analysis, and test design.
• Strong understanding of CPU, GPU, or other high-performance hardware accelerator architectures.
• Familiarity with high-performance computing and distributed programming.
• Experience with CUDA or OpenCL programming is desirable but not mandatory.
• Experience with XLA, TVM, MLIR, LLVM, OpenAI Triton, deep learning models and algorithms, or the design of deep learning frameworks is advantageous.
• Excellent interpersonal skills and the ability to thrive in a dynamic, product-focused team environment.
• Experience with JAX, PyTorch, or TensorFlow; familiarity with CUDA or GPUs; and knowledge of open-source compilers such as XLA, LLVM, MLIR, or TVM will set candidates apart.
• Competitive salaries.
• Comprehensive benefits package.
• Equity opportunities.
• Additional benefits.
Anduril Industries
Sargent & Lundy
Sargent & Lundy
Get handpicked remote jobs straight to your inbox weekly.