Remotery

Senior Deep Learning Engineer, CUDA

Posted Aug 4

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Integrate cutting-edge CUDA features and runtime abstractions into AI frameworks, progressing from proof of concept to performance analysis and eventual production.

• Examine AI workloads and frameworks to uncover lower-level requirements and opportunities for innovation.

• Engage directly with teams focusing on the latest AI models.

• Enhance the AI compiler-runtime interface for solutions involving multiple GPUs and nodes.

• Create fault-tolerant and flexible solutions tailored for large-scale or dynamic AI workloads.

• Shape the core CUDA roadmap for next-generation deep learning frameworks.

• Work collaboratively across various time zones with AI researchers, hardware and software architects, kernel and compiler developers, and CUDA driver specialists.

• Co-design systems and frameworks that boost performance and programmability.

• Develop innovative tools and runtime systems to profile and expedite new deep learning paradigms.

• Write clean, efficient, maintainable code, transitioning prototypes into open-source releases, framework integrations, internal tools, or commercial products.


⛳️ Requirements

• BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.

• Over 8 years of relevant industry experience or comparable academic experience following degree completion.

• Development experience with deep learning frameworks such as PyTorch and JAX, along with inference engines like TRT-LLM, vLLM, and SGLang.

• Proficiency in rapid prototyping and development using Python, C++, CUDA, or similar domain-specific languages.

• Strong comprehension of AI models, parallelism, and/or compiler technologies such as torch.compile.

• Experience in performance benchmarking on AI clusters.

• Familiarity with at least one performance profiling toolchain, like PyTorch Profiler or NVIDIA Nsight Systems.

• Understanding of HPC/AI communication concepts.

• Knowledge of computer system architecture, hardware-software interactions, and operating system principles.

• Flexibility and eagerness to learn new frameworks and tools.

• Capability to work and communicate effectively across diverse teams and time zones.

• Preferred additional expertise in the performance internals of deep learning frameworks and execution graphs.

• Preferred hands-on experience with CUDA, NCCL, MPI, UCX, and distributed machine learning methodologies.

• Preferred knowledge in training, distributed inference, mixture-of-experts, reinforcement learning, or kernel authoring with CUDA, Triton, or cuTe.

• Preferred background in deep learning compilers such as Triton, XLA, and torch.compile.

• Preferred experience in programming compute and communication overlap in distributed runtimes.


🏝️ Benefits

• Equity

• Benefits

People also viewed

NVIDIA8 hours ago

Senior Deep Learning Engineer, Accuracy Evaluation

PL flagPoland, +4 more countriesFull-timeMachine Learning EngineerPLN 375k – PLN 650k/year
ApplyView job
SentiLink9 hours ago

Head of Applied Machine Learning – Application Fraud

US flagUnited States OnlyFull-timeMachine Learning Engineer$210k – $260k/year
ApplyView job
SentiLink9 hours ago

Applied Machine Learning Manager – Application Fraud

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $250k/year
ApplyView job
Leega10 hours ago

Senior AI (Generative, MLOps)

BR flagBrazil OnlyFreelanceMachine Learning Engineer
ApplyView job
Solidgate13 hours ago

Senior Machine Learning Engineer

UA flagUkraine OnlyFull-timeMachine Learning Engineer
ApplyView job
FIT:MATCH.ai13 hours ago

Machine Learning Engineer, 3D Vision

US flagCalifornia, +3 more statesFull-timeMachine Learning Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers