
GPU Kernel Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in New York.
β’ Analyze GPU and accelerator kernels for technical accuracy and comprehensiveness
β’ Examine implementations derived from specifications or reference operators
β’ Evaluate mathematical behavior, implementation issues, unsupported assumptions, and incomplete solutions
β’ Oversee kernel development across CUDA, Triton, NKI, and Pallas (JAX)
β’ Assess framework-specific implementation decisions, execution limitations, translations, and migrations
β’ Compare various kernel implementations for accuracy and technical excellence
β’ Evaluate outputs against reference implementations using absolute, relative, and ULP-based tolerances
β’ Analyze floating-point behavior and precision-related edge cases
β’ Measure performance using Nsight, Nsight Compute, roofline analysis, and framework-native profilers
β’ Review benchmarking methodologies, latency, throughput, utilization, memory behavior, and claimed enhancements
β’ Examine compute- and memory-efficiency optimization strategies, such as tiling, vectorization, parallelization, and workload decomposition
β’ Investigate registers, shared memory, caches, memory access, data locality, bandwidth utilization, bank conflicts, coalescing, and register pressure
β’ Diagnose compilation and runtime failures related to drivers, out-of-memory situations, launch configurations, shape or stride mismatches, and autotuning
β’ Validate task reliability in designated environments and assess debugging techniques
β’ Evaluate kernel translation, lowering, hardware migration, compiler transformations, and intermediate-representation choices
β’ Diagnose kernel implementations using outputs, profiler data, runtime behavior, and source code
β’ Assess operator fusion and the preservation of intended semantics in fused kernels
β’ Evaluate assigned tasks against structured technical criteria and deliver evidence-based written assessments
β’ Differentiate valid implementation alternatives from those with technical flaws
β’ Over 3 years of hands-on experience in developing, optimizing, or validating GPU or accelerator kernels
β’ Practical experience with at least two of the following: CUDA, Triton, NKI, or Pallas (JAX)
β’ Strong grasp of numerical correctness, including absolute, relative, and ULP tolerances
β’ Experience in selecting and validating suitable reference implementations
β’ Extensive performance profiling and benchmarking expertise
β’ Familiarity with Nsight, Nsight Compute, roofline analysis, or similar profiling tools
β’ Strong understanding of common kernel compilation and runtime failure modes
β’ Experience in at least three of the following: kernel generation from specification, framework translation or lowering, hardware-target migration, kernel debugging, performance optimization, operator fusion
β’ Preferred experience in both NVIDIA GPU and custom-accelerator ecosystems
β’ Background in compiler engineering, MLIR, or intermediate-representation lowering is advantageous
β’ Strong understanding of memory-hierarchy optimization is preferred
β’ Contributions to kernel or accelerator libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls are advantageous
β’ Excellent written communication skills and the ability to provide precise technical feedback
β’ Must perform work without utilizing confidential or proprietary information from any employer, client, institution, or other third party
β’ H1-B and STEM OPT support is not available
β’ Opportunity for part-time independent contractor engagement
β’ Fully remote work within the United States
β’ Flexible scheduling based on project requirements
β’ Project duration may be extended, shortened, or concluded depending on project needs and performance
β’ H1-B and STEM OPT support is unavailable for this engagement
Fortinet
Fortinet
Fortinet
Get handpicked remote jobs straight to your inbox weekly.