
AI Research Engineer, Kernel & Inference Optimization
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
• Design and implement cutting-edge model serving architectures that prioritize high throughput, minimal latency, and efficient memory utilization.
• Ensure that inference pipelines operate effectively on resource-constrained devices and edge platforms.
• Set performance benchmarks for latency, token response, and memory usage.
• Construct, execute, and oversee controlled inference tests in both simulated and live production settings.
• Monitor key performance indicators such as latency, throughput, memory consumption, and error rates.
• Document iterative findings and assess results against predefined benchmarks across various platforms.
• Identify and prepare testing datasets and simulation scenarios to tackle low-resource deployment challenges.
• Examine computational efficiency and identify bottlenecks in the serving pipeline.
• Enhance batch processing, network latency, memory management, scalability, and reliability.
• Collaborate with interdisciplinary teams to integrate optimized serving and inference frameworks into production edge and on-device pipelines.
• Establish success metrics and engage in continuous monitoring and iterative improvements.
• A degree in Computer Science or a related discipline.
• Preferably a PhD in NLP, Machine Learning, or a similar field.
• Proven experience in AI research and development, with notable publications in A* conferences.
• Familiarity with Metal Shading Language (MSL).
• Capability to create custom compute shaders from the ground up.
• Demonstrated experience in low-level kernel optimizations and inference enhancements on mobile platforms.
• Tangible improvements in inference latency, throughput, and memory usage for domain-specific applications.
• Extensive knowledge of modern model serving architectures and inference optimization methodologies.
• Strong proficiency in writing GPU kernels for mobile devices, such as smartphones.
• In-depth understanding of model serving frameworks and engines.
• Practical experience in developing and deploying comprehensive inference pipelines on resource-limited devices.
• Ability to apply empirical research to challenges in model serving, including latency optimization, computational bottlenecks, and memory limitations.
• Expertise in designing evaluation frameworks and iterating on optimization techniques.
• Experience with distributed inference systems, including Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism.
• Profound understanding of the mathematics and structure of Diffusion Models and Vision Transformers.
• Knowledge of Pruning, Quantization, Flash Attention, KV Cache, and Speculative Decoding (Eagle).
• Work remotely from any location worldwide.
• Opportunity to collaborate with a diverse global team.
• Engage in innovative projects in fintech, blockchain, AI, and digital finance.
CareSource
Mistral AI
NXP Semiconductors
Get handpicked remote jobs straight to your inbox weekly.