
Research Engineer β Inference
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in United Kingdom.
β’ Implement cutting-edge AI models into production environments
β’ Take ownership of the transition from research milestones to serving infrastructure
β’ Enhance inference performance focusing on latency, throughput, and cost efficiency
β’ Utilize quantization, distillation, KV-cache optimization, batching strategies, and custom kernels
β’ Develop and optimize high-performance serving systems for real-time and streaming workloads
β’ Establish tools and infrastructure that enable researchers to deploy models to production swiftly and securely
β’ Assess and validate the performance characteristics of models
β’ No formal qualifications or degrees necessary
β’ Passion for tackling challenging engineering problems
β’ Capability to showcase work through previous projects, designs, or contributions on GitHub
β’ Experience in deploying and serving ML models in production, particularly for latency-sensitive or real-time applications
β’ Strong engineering expertise in GPU programming and inference optimization, including CUDA, Triton, TensorRT, vLLM, or SGLang
β’ Ability to independently profile, diagnose, and resolve bottlenecks throughout the serving stack
β’ Skill in developing tools to measure the performance of the serving stack
β’ Annual discretionary professional development stipend
β’ Annual discretionary stipend for social travel to connect with colleagues
β’ Annual company offsite
β’ Monthly co-working stipend for employees residing away from a main hub
β’ Option to work from company offices located in London, New York, San Francisco, and Warsaw
Stryker
Anduril Industries
24-MAG
24-MAG
Get handpicked remote jobs straight to your inbox weekly.