
Senior Deep Learning Software Engineer, Inference
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Poland, +2 more states.
β’ Design, construct, and enhance GPU-accelerated software tailored for AI applications.
β’ Develop and sustain high-performance deep learning frameworks such as SGLang and vLLM.
β’ Enhance platforms to streamline the deployment and serving of language models.
β’ Integrate the latest algorithms for public release within SGLang, vLLM, and other deep learning frameworks.
β’ Recognize and drive performance enhancements for cutting-edge LLM and Generative AI models utilizing NVIDIA accelerators.
β’ Utilize open-source tools and plugins, including CUTLASS, OAI Triton, NCCL, and CUDA kernels, to implement and optimize model-serving pipelines.
β’ Optimize, evaluate, and fine-tune deep learning models within the realms of LLM, multimodal, and Generative AI.
β’ Scale deep learning model performance across various architectures and types of NVIDIA accelerators.
β’ Contribute features and code to NVIDIA's inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.
β’ Collaborate with teams across frameworks and NVIDIA libraries to develop inference optimization solutions.
β’ Master's or PhD, or equivalent experience in a pertinent field such as Computer Engineering, Computer Science, EECS, or AI.
β’ A minimum of 5 years of relevant software development experience.
β’ Exceptional C/C++ programming and software design capabilities.
β’ SW Agile skills are advantageous.
β’ Experience with Python is a plus.
β’ Prior experience with training, deploying, or optimizing deep learning model inference in production is beneficial.
β’ A background in performance modeling, profiling, debugging, code optimization, or knowledge of CPU/GPU architectures is a plus.
β’ Experience contributing to deep learning software projects like PyTorch, vLLM, and SGLang is a distinct advantage.
β’ Familiarity with multi-GPU communications, including NCCL and NVSHMEM, is a differentiator.
β’ Experience in building and delivering products to enterprise clients is a differentiator.
β’ GPU programming experience with CUDA, OAI Triton, or CUTLASS is a differentiator.
β’ Highly competitive salaries.
β’ Extensive benefits package.
β’ A work environment that fosters diversity, inclusion, and flexibility.
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.