
Senior Machine Learning Engineer, LLM Inference Optimization
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
• Take charge of optimization tasks for designated model families, customer endpoints, or serving backends.
• Conduct engine comparisons and provide recommendations for effective serving configurations tailored to specific workloads.
• Troubleshoot model quality or performance issues that arise during production rollouts.
• Enhance LLM and VLM endpoints focusing on latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
• Deploy, configure, benchmark, and expand inference engines like vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or comparable systems.
• Develop and implement model-compression workflows, encompassing quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
• Execute or incorporate speculative decoding, draft-model strategies, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
• Create reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
• Collaborate with GPU kernel engineers and platform engineers to identify bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
• Produce clear design documentation, performance reports, rollout strategies, and technical explanations for customers.
• Proficient in Python and PyTorch with strong engineering capabilities.
• Practical experience in deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
• Familiarity with at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or similar internal systems.
• In-depth understanding of transformer inference bottlenecks, including KV cache, attention mechanisms, memory bandwidth, batching, parallelism, and long-context serving.
• Ability to analyze and quantify latency, throughput, quality, utilization, and cost trade-offs.
• Excellent communication skills with the ability to collaborate with research, kernel, infrastructure, product, and customer teams.
• Experience with techniques such as quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or similar methods.
• Familiarity with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration strategies.
• Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
• Knowledge of CUDA or Triton.
• Contributions to open-source projects like vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related initiatives.
• Competitive compensation package.
• Opportunities for career advancement and continuous learning.
• Flexibility and ownership in your role.
• A collaborative and innovative workplace culture.
• Chance to work on significant AI projects.
• An international environment with talented teams.
EVERSANA
i22 Digital Agency
Nagarro
SysMap Solutions
Get handpicked remote jobs straight to your inbox weekly.