
Staff Kernel Optimization Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Create design specifications for new machine learning and linear algebra kernels, mapping them to the Cerebras WSE System utilizing various parallel programming algorithms.
• Develop and troubleshoot a kernel library consisting of highly optimized low-level assembly instructions and routines in a C-like domain-specific language to implement algorithms designed for the Cerebras hardware system.
• Construct and debug high-performance kernel routines using low-level assembly and a custom C-like (CSL) language, optimizing algorithms for the Cerebras hardware environment.
• Utilize mathematical models and analytical methods to evaluate software performance and guide design choices.
• Develop and implement unit and system testing strategies to ensure correct functionality and performance of kernel libraries.
• Investigate emerging trends in Machine Learning applications and assist in evolving the kernel library architecture to tackle the computational challenges posed by state-of-the-art Neural Networks.
• Collaborate with chip and system architects to enhance instruction sets, microarchitecture, and input/output functions of next-generation systems.
• Bachelor’s, Master’s, PhD, or equivalent foreign degree in Computer Science, Computer Engineering, Mathematics, or related disciplines.
• Familiarity with hardware architecture concepts — must be willing to learn the intricacies of a new hardware architecture.
• Proficient in C++ and Python programming languages.
• Solid understanding of library and/or API development best practices.
• Strong debugging capabilities and experience with complex software stack debugging.
• Experience in kernel development and/or testing (preferred).
• Knowledge of parallel algorithms and distributed memory systems (preferred).
• Experience with programming accelerators such as GPUs and FPGAs (preferred).
• Familiarity with Machine Learning neural networks and frameworks like TensorFlow and PyTorch (preferred).
• Understanding of HPC kernels and their optimization (preferred).
• Create a revolutionary AI platform that transcends the limitations of traditional GPU technology.
• Publish and contribute to open-source cutting-edge AI research.
• Work on one of the fastest AI supercomputers worldwide.
• Experience job stability with the dynamic energy of a startup.
• Enjoy a straightforward, non-corporate work culture that honors individual beliefs.
Fortinet
Fortinet
Get handpicked remote jobs straight to your inbox weekly.