Senior Machine Learning Engineer, LLM Inference Optimization

Posted 2 days ago

This is a fully remote position, open to applicants in Europe.

📋 Description

• Take charge of optimization tasks for designated model families, customer endpoints, or serving backends.

• Conduct engine comparisons and provide recommendations for effective serving configurations tailored to specific workloads.

• Troubleshoot model quality or performance issues that arise during production rollouts.

• Enhance LLM and VLM endpoints focusing on latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.

• Deploy, configure, benchmark, and expand inference engines like vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or comparable systems.

• Develop and implement model-compression workflows, encompassing quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.

• Execute or incorporate speculative decoding, draft-model strategies, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.

• Create reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.

• Collaborate with GPU kernel engineers and platform engineers to identify bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.

• Produce clear design documentation, performance reports, rollout strategies, and technical explanations for customers.


⛳️ Requirements

• Proficient in Python and PyTorch with strong engineering capabilities.

• Practical experience in deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.

• Familiarity with at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or similar internal systems.

• In-depth understanding of transformer inference bottlenecks, including KV cache, attention mechanisms, memory bandwidth, batching, parallelism, and long-context serving.

• Ability to analyze and quantify latency, throughput, quality, utilization, and cost trade-offs.

• Excellent communication skills with the ability to collaborate with research, kernel, infrastructure, product, and customer teams.

• Experience with techniques such as quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or similar methods.

• Familiarity with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration strategies.

• Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.

• Knowledge of CUDA or Triton.

• Contributions to open-source projects like vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related initiatives.


🏝️ Benefits

• Competitive compensation package.

• Opportunities for career advancement and continuous learning.

• Flexibility and ownership in your role.

• A collaborative and innovative workplace culture.

• Chance to work on significant AI projects.

• An international environment with talented teams.

People also viewed

EVERSANA1 day ago

Conversational AI Engineer

US flagKansas OnlyFull-timeLLM Engineer$100k – $125k/year
ApplyView job
i22 Digital Agency3 days ago

Senior AI Software Engineer – LLM Applications

DE flagGermany OnlyFull-timeLLM Engineer
ApplyView job
Nagarro3 days ago

Associate Staff Engineer – Conversational AI Developer

US flagPennsylvania OnlyFull-timeLLM Engineer
ApplyView job
SysMap Solutions4 days ago

Senior Generative AI Engineer

BR flagBrazil OnlyFull-timeLLM Engineer
ApplyView job
Azumo4 days ago

AI Engineer – Generative AI, Agents

AR flagArgentina, +5 more countriesFull-timeLLM Engineer
ApplyView job
Nagarro5 days ago

Senior Staff Engineer, Generative AI

US flagPennsylvania OnlyFull-timeLLM Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers