Remotery

CUDA Engineer

Posted Jul 19

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Develop and enhance custom CUDA kernels for fundamental transformer inference tasks.

• Analyze kernels to identify and resolve bottlenecks related to occupancy, memory throughput, and warp divergence.

• Implement kernel fusion techniques to minimize memory round-trips and reduce launch overhead across inference workflows.

• Optimize memory access patterns and oversee the memory hierarchy for optimal bandwidth utilization.

• Create quantization-aware kernels and mixed-precision arithmetic to decrease latency and memory usage.

• Design and adjust caching mechanisms for efficient autoregressive decoding.

• Fine-tune kernel launch configurations tailored for specific GPU architectures.

• Evaluate kernels against existing benchmarks to achieve tangible improvements in throughput and latency.

• Develop tests for CUDA code to identify performance and correctness regressions.

• Maintain internal CUDA libraries while contributing to team coding standards and documentation.


⛳️ Requirements

• A minimum of 4 years of experience writing production CUDA code, with a proven record of delivering performance-critical kernels.

• Profound knowledge of GPU microarchitecture, including warps, occupancy, register pressure, and memory hierarchy.

• Strong skills in CUDA C++, particularly in streams and asynchronous execution.

• Practical experience in profiling to distinguish compute-bound from memory-bound bottlenecks.

• Familiarity with kernel fusion, memory coalescing, and strategies to avoid warp divergence.

• Proven experience in creating quantized and mixed-precision kernels.

• A solid understanding of parallel algorithm design and the trade-offs of numerical precision.

• Nice to Have

• Experience with transformer or attention-style kernels, or autoregressive decoding.

• Experience in developing high-performance GPU libraries from the ground up.

• Background in HPC or other fields that prioritize latency-critical performance engineering.

• Exposure to optimization at the kernel level for multi-GPU or multi-node systems.

• Comfort in reading PTX/SASS to assess kernel efficiency.


🏝️ Benefits

• Competitive salary along with an equity sign-on bonus.

• Biannual bonus program.

• Fully covered technology expenses tailored to your requirements.

• Breakfast and dinner allowances for employees working in the office.

People also viewed

ABBJul 26

Application and Promotion Engineer

BR flagBrazil OnlyFull-timeEngineer
ApplyView job
3M ConsultancyJul 26

Kafka Engineer, IRS MBI Clearance

US flagUnited States OnlyFull-timeEngineer
ApplyView job
mpathicJul 26

Tier 2 Networking Engineer

US flagUnited States OnlyFull-timeEngineer
ApplyView job
Jones Lang LaSalle Americas, Inc.Jul 26

Forward Deploy Engineer

US flagCalifornia, +1 more stateFull-timeEngineer$220k – $320k/year
ApplyView job
PragmatikeJul 25

Industrial Process Engineer

CL flagChile OnlyFull-timeEngineer
ApplyView job
Interlaced.ioJul 25

Technical Project Engineer

US flagUnited States OnlyFreelanceEngineer$65k – $70k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers