
GPU Kernel Engineer β CUDA, Triton, Accelerator Performance
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in Argentina.
β’ Analyze GPU and accelerator kernel implementations for accuracy
β’ Validate outputs against established reference implementations
β’ Assess numerical tolerance levels
β’ Examine kernel benchmarks to ensure fair comparisons
β’ Detect performance bottlenecks and identify optimization possibilities
β’ Evaluate if performance goals are attainable given hardware limitations
β’ Review kernel translations and transitions across hardware
β’ Identify issues related to compilation, drivers, memory, shapes, and runtime
β’ Determine if technical tasks are genuinely challenging or misconfigured
β’ Offer clear and actionable technical insights
β’ Implement and troubleshoot kernels
β’ Enhance the performance of CUDA and Triton kernels
β’ Translate between different kernel frameworks
β’ Conduct hardware migrations and operator fusion
β’ Profile and benchmark system performance
β’ Confirm numerical accuracy
β’ Troubleshoot compilation and runtime challenges
β’ Optimize memory hierarchy and AI workload performance at the kernel level
β’ Over 3 years of practical experience in developing, optimizing, or debugging GPU or accelerator kernels
β’ Significant experience with at least two of the following: CUDA; Triton; NKI / AWS Neuron; Pallas / JAX
β’ Strong grasp of GPU performance optimization techniques
β’ Familiarity with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-specific profilers
β’ Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts
β’ Solid comprehension of floating-point numerical accuracy and tolerance levels
β’ Experience in debugging kernel compilation and runtime problems
β’ Capability to differentiate between software defects, environmental issues, and genuine optimization challenges
β’ Experience in developing kernels from technical specifications, translating kernels across frameworks, migrating kernels to different hardware platforms, debugging faulty implementations, optimizing kernel efficiency, and combining multiple operations into optimized kernels
β’ Exposure to both NVIDIA GPU and custom accelerator ecosystems (preferred)
β’ Familiarity with AWS Trainium, TPU, JAX, or other accelerators (preferred)
β’ Experience in compiler engineering (preferred)
β’ Knowledge of MLIR, XLA, or lowering intermediate representations (preferred)
β’ Contributions to GPU or ML kernel libraries (preferred)
β’ Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls (preferred)
β’ Background in AI model evaluation, RLHF, or technical benchmark development (preferred)
β’ Compensation of $65 per hour
β’ Part-time, project-based consulting opportunity
β’ Flexible remote work arrangement
DriveNets
ServiceTitan
SEH
Get handpicked remote jobs straight to your inbox weekly.