Remotery

AI Inference Engineer

Posted 2 hours ago

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Establish Fuse's strategy and architecture for inference serving based on foundational principles.

• Create and develop the serving stack, including request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference tasks.

• Lead the model-level optimization strategy for serving, determining the application of techniques such as quantization, distillation, and speculative decoding to enhance throughput and reduce cost per token, in collaboration with CUDA/GPU engineers.

• Make essential software architecture decisions regarding serving frameworks and orchestration (e.g., vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or similar technologies).

• Convert throughput, latency, and uptime obligations into actionable technical specifications and serving capacity plans.

• Serve as the primary technical authority on inference performance and reliability.

• Collaborate closely with CUDA and GPU engineering teams to ensure seamless integration of custom kernels and hardware performance work into the serving layer.

• Establish the standards, tools, and benchmarks that this function will utilize as it expands.


⛳️ Requirements

• A minimum of 4 years of experience in building or managing large-scale inference serving systems, or equivalent substantial project/industry experience.

• Extensive hands-on experience with inference serving frameworks and the optimization techniques associated with them (batching, KV-cache management, quantization, speculative decoding).

• Strong systems thinking ability, capable of understanding the complete process from incoming requests to served responses across a large cluster.

• Proficient in collaborating directly with GPU/CUDA engineers to incorporate low-level performance optimizations into a serving system.

• Proven history of making critical architectural decisions and taking ownership of their outcomes.

• Ability to operate independently without predefined guidelines - this is a pioneering role that will shape a new function around architecture that is still in its early stages, rather than joining an established team.

• **Nice to Have**

• Familiarity with Triton or custom ML inference/training frameworks.

• Experience with autoscaling or capacity planning for large-scale inference tasks.

• Exposure to multi-tenant serving or SLA-driven infrastructure.

• Background in a hyperscaler, frontier AI lab, or large-scale distributed inference systems.

• Knowledge of Kubernetes/Slurm for cluster orchestration.

• Interest or experience in energy markets, grid systems, or sustainable computing.


🏝️ Benefits

• Competitive salary along with an equity sign-on bonus.

• Biannual bonus program.

• Fully funded technology to meet your requirements.

• Breakfast and dinner allowances for office-based employees.

People also viewed

Zeus2 hours ago

AI Engineer

US flagSouth Carolina OnlyFull-timeAI Engineer
ApplyView job
Udacity Marketing2 hours ago

AI Engineer Technical Mentor – Independent Contractor

North AmericaFreelanceAI Engineer
ApplyView job
FlexiSAF Edusoft Limited2 hours ago

AI Engineering Intern

NG flagNigeria OnlyFull-timeAI Engineer
ApplyView job
BJAK2 hours ago

Mobile Developer – AI Neobank App

TH flagThailand OnlyFull-timeAI Engineer
ApplyView job
Autodesk2 hours ago

Machine Learning Engineer, 3D Geometry, Multimodal AI

CA flagCanada OnlyFull-timeAI Engineer$123k – $180.4k/year
ApplyView job
Kainos2 hours ago

AI Engineer – Workday Products

PL flagPoland OnlyFull-timeAI Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers