GPU Kernel Engineer – CUDA, Triton, Accelerator Performance

atAnyone AIRemoteAR flagArgentinaFreelanceEngineerMid-levelSenior$65/hour

Posted 3 hours ago

This is a fully remote position, open to applicants in Argentina.

πŸ“‹ Description

β€’ Analyze GPU and accelerator kernel implementations for accuracy

β€’ Validate outputs against established reference implementations

β€’ Assess numerical tolerance levels

β€’ Examine kernel benchmarks to ensure fair comparisons

β€’ Detect performance bottlenecks and identify optimization possibilities

β€’ Evaluate if performance goals are attainable given hardware limitations

β€’ Review kernel translations and transitions across hardware

β€’ Identify issues related to compilation, drivers, memory, shapes, and runtime

β€’ Determine if technical tasks are genuinely challenging or misconfigured

β€’ Offer clear and actionable technical insights

β€’ Implement and troubleshoot kernels

β€’ Enhance the performance of CUDA and Triton kernels

β€’ Translate between different kernel frameworks

β€’ Conduct hardware migrations and operator fusion

β€’ Profile and benchmark system performance

β€’ Confirm numerical accuracy

β€’ Troubleshoot compilation and runtime challenges

β€’ Optimize memory hierarchy and AI workload performance at the kernel level


⛳️ Requirements

β€’ Over 3 years of practical experience in developing, optimizing, or debugging GPU or accelerator kernels

β€’ Significant experience with at least two of the following: CUDA; Triton; NKI / AWS Neuron; Pallas / JAX

β€’ Strong grasp of GPU performance optimization techniques

β€’ Familiarity with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-specific profilers

β€’ Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts

β€’ Solid comprehension of floating-point numerical accuracy and tolerance levels

β€’ Experience in debugging kernel compilation and runtime problems

β€’ Capability to differentiate between software defects, environmental issues, and genuine optimization challenges

β€’ Experience in developing kernels from technical specifications, translating kernels across frameworks, migrating kernels to different hardware platforms, debugging faulty implementations, optimizing kernel efficiency, and combining multiple operations into optimized kernels

β€’ Exposure to both NVIDIA GPU and custom accelerator ecosystems (preferred)

β€’ Familiarity with AWS Trainium, TPU, JAX, or other accelerators (preferred)

β€’ Experience in compiler engineering (preferred)

β€’ Knowledge of MLIR, XLA, or lowering intermediate representations (preferred)

β€’ Contributions to GPU or ML kernel libraries (preferred)

β€’ Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls (preferred)

β€’ Background in AI model evaluation, RLHF, or technical benchmark development (preferred)


🏝️ Benefits

β€’ Compensation of $65 per hour

β€’ Part-time, project-based consulting opportunity

β€’ Flexible remote work arrangement

People also viewed

DriveNets4 hours ago

Senior Signal Integrity Engineer

TW flagTaiwan OnlyFull-timeEngineer
ApplyView job
Marvik5 hours ago

Embedded Software, Controls Engineer

UY flagUruguay OnlyFull-timeEngineer
ApplyView job
ServiceTitan5 hours ago

Project Engineer I

US flagCalifornia, +8 more statesFull-timeEngineer$94.7k – $152k/year
ApplyView job
SEH5 hours ago

Wastewater Treatment Engineer – Entry Level

US flagWisconsin OnlyFull-timeEngineer$65k – $75k/year
ApplyView job
AECOM5 hours ago

Senior/Principal Process Engineer

US flagNew Mexico OnlyFull-timeEngineer
ApplyView job
ENTRUST Solutions Group6 hours ago

Senior Substation Project Engineer

US flagCalifornia, +1 more stateFull-timeEngineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers