Senior Deep Learning Software Engineer, Inference

Posted Aug 11

This is a fully remote position, open to applicants in Poland, +2 more countries.

πŸ“‹ Description

β€’ Design, construct, and enhance GPU-accelerated software tailored for AI applications.

β€’ Develop and sustain high-performance deep learning frameworks such as SGLang and vLLM.

β€’ Enhance platforms to streamline the deployment and serving of language models.

β€’ Integrate the latest algorithms for public release within SGLang, vLLM, and other deep learning frameworks.

β€’ Recognize and drive performance enhancements for cutting-edge LLM and Generative AI models utilizing NVIDIA accelerators.

β€’ Utilize open-source tools and plugins, including CUTLASS, OAI Triton, NCCL, and CUDA kernels, to implement and optimize model-serving pipelines.

β€’ Optimize, evaluate, and fine-tune deep learning models within the realms of LLM, multimodal, and Generative AI.

β€’ Scale deep learning model performance across various architectures and types of NVIDIA accelerators.

β€’ Contribute features and code to NVIDIA's inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.

β€’ Collaborate with teams across frameworks and NVIDIA libraries to develop inference optimization solutions.


⛳️ Requirements

β€’ Master's or PhD, or equivalent experience in a pertinent field such as Computer Engineering, Computer Science, EECS, or AI.

β€’ A minimum of 5 years of relevant software development experience.

β€’ Exceptional C/C++ programming and software design capabilities.

β€’ SW Agile skills are advantageous.

β€’ Experience with Python is a plus.

β€’ Prior experience with training, deploying, or optimizing deep learning model inference in production is beneficial.

β€’ A background in performance modeling, profiling, debugging, code optimization, or knowledge of CPU/GPU architectures is a plus.

β€’ Experience contributing to deep learning software projects like PyTorch, vLLM, and SGLang is a distinct advantage.

β€’ Familiarity with multi-GPU communications, including NCCL and NVSHMEM, is a differentiator.

β€’ Experience in building and delivering products to enterprise clients is a differentiator.

β€’ GPU programming experience with CUDA, OAI Triton, or CUTLASS is a differentiator.


🏝️ Benefits

β€’ Highly competitive salaries.

β€’ Extensive benefits package.

β€’ A work environment that fosters diversity, inclusion, and flexibility.

People also viewed

Innovecs1 day ago

Senior Full-stack Engineer, Angular, Node.js

EuropeFull-timeFull-stack Engineer
ApplyView job
saas.group1 day ago

Senior Software Engineer

Anywhere in the WorldFull-timeFull-stack Engineer
ApplyView job
Dremio1 day ago

Software Engineer – Query Execution

PT flagPortugal OnlyFull-timeFull-stack Engineer
ApplyView job
Speechify1 day ago

Tech Lead, Android Core Product

PL flagPoland OnlyFull-timeFull-stack Engineer$30k – $110k/year
ApplyView job
DDN1 day ago

Staff Software Engineer

US flagCalifornia OnlyFull-timeFull-stack Engineer$150k – $250k/year
ApplyView job
DecisionPoint Corporation1 day ago

Full Stack Developer – React/Java

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers