AWS Trainium Kernel Engineer

at24-MAGRemoteUS flagNew YorkFreelanceEngineerJuniorMid-level$60 – $80/hour

Posted 16 hours ago

This is a fully remote position, open to applicants in New York.

πŸ“‹ Description

β€’ Evaluate kernels developed utilizing the Neuron Kernel Interface (NKI)

β€’ Determine if implementations effectively target AWS Trainium and Inferentia2 hardware

β€’ Examine low-level computation patterns for technical accuracy

β€’ Review the migration of existing CUDA kernels to NKI

β€’ Assess tile decomposition, computational strategies, and partitioning choices

β€’ Analyze the use of SBUF, PSUM, and HBM across Trainium memory hierarchies

β€’ Evaluate DMA orchestration, data movement, sequencing, stalls, and synchronization

β€’ Review Trainium-specific optimization strategies and identify performance bottlenecks

β€’ Assess NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth behavior

β€’ Evaluate numerical consistency between GPU and Trainium implementations

β€’ Review precision choices, mixed-precision behavior, numerical stability, and supported formats

β€’ Evaluate implementations using the AWS Neuron SDK, compiler, and NKI kernel libraries

β€’ Analyze benchmark results for Trn1 or Trn2 workloads and evaluate experimental methodology

β€’ Assess assigned kernel-development tasks against established technical criteria

β€’ Provide clear written, rubric-based technical feedback supported by implementation, profiling, or numerical evidence


⛳️ Requirements

β€’ 2+ years of hands-on experience in developing or optimizing kernels using the Neuron Kernel Interface (NKI)

β€’ Professional experience targeting AWS Trainium or Inferentia2 hardware

β€’ Strong understanding of tile-based computation

β€’ Deep familiarity with SBUF, PSUM, and HBM memory management

β€’ Strong knowledge of partition-dimension constraints and DMA orchestration

β€’ Proven experience evaluating or conducting CUDA-to-NKI migrations

β€’ Familiarity with Trainium-specific performance profiling

β€’ Experience in assessing NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks

β€’ Strong understanding of cross-platform numerical correctness and mixed-precision behavior

β€’ Excellent written communication skills and ability to provide precise technical feedback

β€’ Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred

β€’ Prior experience in CUDA or Triton kernel development is advantageous

β€’ Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred

β€’ Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous

β€’ H1-B and STEM OPT support is not available for this engagement


🏝️ Benefits

β€’ Part-time independent contractor engagement

β€’ Fully remote position within the United States

β€’ Flexible scheduling based on project requirements

β€’ Projects may be extended, shortened, or concluded based on project needs and performance

People also viewed

Ford Motor Company17 hours ago

Cloud Platform Virtualization Engineer

US flagMichigan OnlyFull-timeEngineer$85.4k – $144.9k/year
ApplyView job
Crusoe17 hours ago

Senior Facilities Engineer

US flagTexas OnlyFull-timeEngineer$90k – $110k/year
ApplyView job
Greenhouse Software17 hours ago

Forward Deployed Engineer

US flagIllinois, +4 more statesFull-timeEngineer$130k – $170k/year
ApplyView job
Interview Pen17 hours ago

Interview Engineer

TH flagThailand OnlyFreelanceEngineer
ApplyView job
Thermo Systems18 hours ago

Control Systems Project Engineer II

US flagOhio OnlyFull-timeEngineer
ApplyView job
Leidos19 hours ago

Lead Transmission Line Engineer

US flagOhio, +1 more stateFull-timeEngineer$92.3k – $166.9k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers