Senior Solutions Architect – Large Scale Neural Networks Inference

Posted Sep 8

This is a fully remote position, open to applicants in France, +4 more countries.

πŸ“‹ Description

β€’ Spearhead the inference strategy for a portfolio of EMEA AI Natives customers, overseeing engagements from the initial proof of concept through to production-scale implementations.

β€’ Recognize inference challenges within customer deployments, including latency, efficiency, cost per token, memory usage, and low-latency networking.

β€’ Design and enhance high-performance inference pipelines utilizing NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other inference backends.

β€’ Enhance GPU utilization and improve the efficiency of AI clusters.

β€’ Convert customer insights and deployment trends into actionable product feedback.

β€’ Formulate the development roadmap for the NVIDIA stack, encompassing Dynamo, TensorRT-LLM, and NIM.

β€’ Establish the technical direction for AI inference throughout the EMEA region.

β€’ Coordinate between NVIDIA and customer organization stakeholders to shape strategic technology decisions for next-generation AI inference at scale.


⛳️ Requirements

β€’ MS or PhD in Computer Science, Engineering, or equivalent experience in the domain.

β€’ Over 8 years of experience in AI/ML infrastructure, with substantial expertise in LLM/VLM inference optimization and large-scale production deployment.

β€’ In-depth knowledge of transformer inference acceleration techniques, including quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models.

β€’ Familiarity with GPU memory hierarchies and low-latency networking and their impact on inference performance.

β€’ Demonstrated ability to lead technical initiatives successfully.

β€’ Strong communication skills, adept at engaging with research scientists, infrastructure engineers, and executive team members.

β€’ Experience with NVIDIA's inference stack, including TensorRT-LLM, Triton Inference Server, NIM, and NVIDIA Dynamo.

β€’ Proficient in GPU orchestration on Kubernetes.

β€’ Experience managing inference at scale within a leading AI lab or a hyperscale inference team.

β€’ Contributions to open-source inference projects such as vLLM, SGLang, KServe, or NVIDIA Dynamo.


🏝️ Benefits

β€’ Highly competitive salaries.

β€’ Comprehensive benefits package.

People also viewed

U.S. Bank1 day ago

Senior Solutions Consultant – Technical Sales Specialist

US flagArizona, +6 more statesFull-timeSolutions Engineer$141.2k – $172.6k/year
ApplyView job
First Quality1 day ago

Solution Architect – Manufacturing Systems

US flagOhio, +2 more statesFull-timeSolutions Engineer
ApplyView job
Arista Networks1 day ago

Technical Solutions Engineer

IE flagIreland OnlyFull-timeSolutions Engineer
ApplyView job
Miovision1 day ago

Technical Solution Engineer

US flagUnited States OnlyFull-timeSolutions Engineer
ApplyView job
CDW1 day ago

Principal Solution Architect – Hybrid Infrastructure

US flagUnited States OnlyFull-timeSolutions Engineer$151.4k – $211.9k/year
ApplyView job
Blue Yonder1 day ago

Enterprise Solution Architect – Manufacturing Planning

US flagUnited States OnlyFull-timeSolutions Engineer$85.2k – $164k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers