
AWS Trainium Kernel Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in New York.
β’ Evaluate kernels developed utilizing the Neuron Kernel Interface (NKI)
β’ Determine if implementations effectively target AWS Trainium and Inferentia2 hardware
β’ Examine low-level computation patterns for technical accuracy
β’ Review the migration of existing CUDA kernels to NKI
β’ Assess tile decomposition, computational strategies, and partitioning choices
β’ Analyze the use of SBUF, PSUM, and HBM across Trainium memory hierarchies
β’ Evaluate DMA orchestration, data movement, sequencing, stalls, and synchronization
β’ Review Trainium-specific optimization strategies and identify performance bottlenecks
β’ Assess NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth behavior
β’ Evaluate numerical consistency between GPU and Trainium implementations
β’ Review precision choices, mixed-precision behavior, numerical stability, and supported formats
β’ Evaluate implementations using the AWS Neuron SDK, compiler, and NKI kernel libraries
β’ Analyze benchmark results for Trn1 or Trn2 workloads and evaluate experimental methodology
β’ Assess assigned kernel-development tasks against established technical criteria
β’ Provide clear written, rubric-based technical feedback supported by implementation, profiling, or numerical evidence
β’ 2+ years of hands-on experience in developing or optimizing kernels using the Neuron Kernel Interface (NKI)
β’ Professional experience targeting AWS Trainium or Inferentia2 hardware
β’ Strong understanding of tile-based computation
β’ Deep familiarity with SBUF, PSUM, and HBM memory management
β’ Strong knowledge of partition-dimension constraints and DMA orchestration
β’ Proven experience evaluating or conducting CUDA-to-NKI migrations
β’ Familiarity with Trainium-specific performance profiling
β’ Experience in assessing NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks
β’ Strong understanding of cross-platform numerical correctness and mixed-precision behavior
β’ Excellent written communication skills and ability to provide precise technical feedback
β’ Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
β’ Prior experience in CUDA or Triton kernel development is advantageous
β’ Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred
β’ Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
β’ H1-B and STEM OPT support is not available for this engagement
β’ Part-time independent contractor engagement
β’ Fully remote position within the United States
β’ Flexible scheduling based on project requirements
β’ Projects may be extended, shortened, or concluded based on project needs and performance
Ford Motor Company
Crusoe
Greenhouse Software
Get handpicked remote jobs straight to your inbox weekly.