
CUDA Engineer
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in United Kingdom.
• Develop and enhance custom CUDA kernels for fundamental transformer inference tasks.
• Analyze kernels to identify and resolve bottlenecks related to occupancy, memory throughput, and warp divergence.
• Implement kernel fusion techniques to minimize memory round-trips and reduce launch overhead across inference workflows.
• Optimize memory access patterns and oversee the memory hierarchy for optimal bandwidth utilization.
• Create quantization-aware kernels and mixed-precision arithmetic to decrease latency and memory usage.
• Design and adjust caching mechanisms for efficient autoregressive decoding.
• Fine-tune kernel launch configurations tailored for specific GPU architectures.
• Evaluate kernels against existing benchmarks to achieve tangible improvements in throughput and latency.
• Develop tests for CUDA code to identify performance and correctness regressions.
• Maintain internal CUDA libraries while contributing to team coding standards and documentation.
• A minimum of 4 years of experience writing production CUDA code, with a proven record of delivering performance-critical kernels.
• Profound knowledge of GPU microarchitecture, including warps, occupancy, register pressure, and memory hierarchy.
• Strong skills in CUDA C++, particularly in streams and asynchronous execution.
• Practical experience in profiling to distinguish compute-bound from memory-bound bottlenecks.
• Familiarity with kernel fusion, memory coalescing, and strategies to avoid warp divergence.
• Proven experience in creating quantized and mixed-precision kernels.
• A solid understanding of parallel algorithm design and the trade-offs of numerical precision.
• Nice to Have
• Experience with transformer or attention-style kernels, or autoregressive decoding.
• Experience in developing high-performance GPU libraries from the ground up.
• Background in HPC or other fields that prioritize latency-critical performance engineering.
• Exposure to optimization at the kernel level for multi-GPU or multi-node systems.
• Comfort in reading PTX/SASS to assess kernel efficiency.
• Competitive salary along with an equity sign-on bonus.
• Biannual bonus program.
• Fully covered technology expenses tailored to your requirements.
• Breakfast and dinner allowances for employees working in the office.
3M Consultancy
Get handpicked remote jobs straight to your inbox weekly.