Remotery

Senior Deep Learning Software Engineer, Inference

Posted 3 days ago

This is a fully remote position, open to applicants in Poland, +2 more states.

πŸ“‹ Description

β€’ Design, construct, and enhance GPU-accelerated software tailored for AI applications.

β€’ Develop and sustain high-performance deep learning frameworks such as SGLang and vLLM.

β€’ Enhance platforms to streamline the deployment and serving of language models.

β€’ Integrate the latest algorithms for public release within SGLang, vLLM, and other deep learning frameworks.

β€’ Recognize and drive performance enhancements for cutting-edge LLM and Generative AI models utilizing NVIDIA accelerators.

β€’ Utilize open-source tools and plugins, including CUTLASS, OAI Triton, NCCL, and CUDA kernels, to implement and optimize model-serving pipelines.

β€’ Optimize, evaluate, and fine-tune deep learning models within the realms of LLM, multimodal, and Generative AI.

β€’ Scale deep learning model performance across various architectures and types of NVIDIA accelerators.

β€’ Contribute features and code to NVIDIA's inference libraries, vLLM, SGLang, FlashInfer, and LLM software solutions.

β€’ Collaborate with teams across frameworks and NVIDIA libraries to develop inference optimization solutions.


⛳️ Requirements

β€’ Master's or PhD, or equivalent experience in a pertinent field such as Computer Engineering, Computer Science, EECS, or AI.

β€’ A minimum of 5 years of relevant software development experience.

β€’ Exceptional C/C++ programming and software design capabilities.

β€’ SW Agile skills are advantageous.

β€’ Experience with Python is a plus.

β€’ Prior experience with training, deploying, or optimizing deep learning model inference in production is beneficial.

β€’ A background in performance modeling, profiling, debugging, code optimization, or knowledge of CPU/GPU architectures is a plus.

β€’ Experience contributing to deep learning software projects like PyTorch, vLLM, and SGLang is a distinct advantage.

β€’ Familiarity with multi-GPU communications, including NCCL and NVSHMEM, is a differentiator.

β€’ Experience in building and delivering products to enterprise clients is a differentiator.

β€’ GPU programming experience with CUDA, OAI Triton, or CUTLASS is a differentiator.


🏝️ Benefits

β€’ Highly competitive salaries.

β€’ Extensive benefits package.

β€’ A work environment that fosters diversity, inclusion, and flexibility.

People also viewed

Cloudera6 hours ago

Staff Software Engineer, Flink/Streaming Analytics

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job
Stellar Cyber7 hours ago

Senior / Staff Software Engineer – Parser Team

US flagUnited States OnlyFull-timeFull-stack Engineer$150k – $200k/year
ApplyView job
Pragmatike7 hours ago

Senior Founding Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike7 hours ago

Staff Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike7 hours ago

Founding Product Engineer – YCombinator Experience

US flagCalifornia OnlyFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike7 hours ago

Lead Product Engineer

US flagCalifornia, +2 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers