
AI Inference Engineer
Posted 2 hours ago

Posted 2 hours ago
This is a fully remote position, open to applicants in United Kingdom.
• Establish Fuse's strategy and architecture for inference serving based on foundational principles.
• Create and develop the serving stack, including request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference tasks.
• Lead the model-level optimization strategy for serving, determining the application of techniques such as quantization, distillation, and speculative decoding to enhance throughput and reduce cost per token, in collaboration with CUDA/GPU engineers.
• Make essential software architecture decisions regarding serving frameworks and orchestration (e.g., vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or similar technologies).
• Convert throughput, latency, and uptime obligations into actionable technical specifications and serving capacity plans.
• Serve as the primary technical authority on inference performance and reliability.
• Collaborate closely with CUDA and GPU engineering teams to ensure seamless integration of custom kernels and hardware performance work into the serving layer.
• Establish the standards, tools, and benchmarks that this function will utilize as it expands.
• A minimum of 4 years of experience in building or managing large-scale inference serving systems, or equivalent substantial project/industry experience.
• Extensive hands-on experience with inference serving frameworks and the optimization techniques associated with them (batching, KV-cache management, quantization, speculative decoding).
• Strong systems thinking ability, capable of understanding the complete process from incoming requests to served responses across a large cluster.
• Proficient in collaborating directly with GPU/CUDA engineers to incorporate low-level performance optimizations into a serving system.
• Proven history of making critical architectural decisions and taking ownership of their outcomes.
• Ability to operate independently without predefined guidelines - this is a pioneering role that will shape a new function around architecture that is still in its early stages, rather than joining an established team.
• **Nice to Have**
• Familiarity with Triton or custom ML inference/training frameworks.
• Experience with autoscaling or capacity planning for large-scale inference tasks.
• Exposure to multi-tenant serving or SLA-driven infrastructure.
• Background in a hyperscaler, frontier AI lab, or large-scale distributed inference systems.
• Knowledge of Kubernetes/Slurm for cluster orchestration.
• Interest or experience in energy markets, grid systems, or sustainable computing.
• Competitive salary along with an equity sign-on bonus.
• Biannual bonus program.
• Fully funded technology to meet your requirements.
• Breakfast and dinner allowances for office-based employees.
Udacity Marketing
FlexiSAF Edusoft Limited
Get handpicked remote jobs straight to your inbox weekly.