
Research Scientist β Performance Optimization
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
β’ Analyze and enhance GPU, CPU, and accelerator code to ensure optimal utilization and minimal latency.
β’ Create high-performance code using PyTorch, Triton, and CUDA, including developing custom operations when necessary.
β’ Develop fused kernels and utilize tensor cores along with modern hardware capabilities across various platforms.
β’ Enhance model architectures and implementations for distributed multi-node production deployment.
β’ Create tools for performance monitoring, analysis, and automation.
β’ Investigate and apply state-of-the-art optimization techniques for transformer models.
β’ Profile existing training and inference processes within the first 30 days.
β’ Deliver and validate a kernel or architecture optimization during the 30 to 60-day period.
β’ Establish monitoring and automation to mitigate performance regressions in the 60 to 90 days timeframe.
β’ Advanced expertise in Triton/CUDA programming and GPU optimization.
β’ Proficient in PyTorch, particularly in kernel development and custom operations.
β’ Familiarity with profiling tools, such as NVIDIA Nsight, torch profiler, and custom tooling.
β’ Comprehensive knowledge of transformer architectures and attention mechanisms.
β’ Experience with compilers and exporters like torch.compile, TensorRT, ONNX, or XLA (preferred).
β’ Background in optimizing inference workloads for both latency and throughput (preferred).
β’ Understanding of Triton compiler and kernel fusion techniques (preferred).
β’ Knowledge of warp-level intrinsics and advanced CUDA optimization techniques (preferred).
β’ Equal opportunity employer.
β’ Participation in a voluntary diversity and inclusion survey; opting out will not impact the job application.
Praxis
Praxis
DECA
Get handpicked remote jobs straight to your inbox weekly.