Remotery

Staff Software Engineer, AI Inference

atSylloRemoteUS flagUnited StatesFull-timeAI EngineerLead$190k – $230k/year

Posted Jul 29

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee the design and development of our production inference platform.

• Establish the technical roadmap for inference infrastructure, model serving, and runtime optimization.

• Create and manage scalable, cost-efficient systems for deploying large language models in a production environment.

• Assess and incorporate contemporary inference technologies, frameworks, and serving runtimes.

• Enhance latency, throughput, GPU utilization, memory efficiency, and overall infrastructure costs.

• Construct systems for model deployment, traffic management, autoscaling, scheduling, observability, and operational excellence.

• Collaborate with ML engineers to transition new models and inference techniques into production.

• Develop benchmarking methodologies to assess new models, runtimes, and hardware.

• Make critical architectural decisions regarding when to develop internally versus utilizing open-source or commercial solutions.

• Guide engineers as the team expands and assist in establishing engineering best practices for AI infrastructure.


⛳️ Requirements

• Extensive experience in designing and managing production AI inference systems.

• Background in building or leading production LLM serving infrastructure.

• In-depth experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or similar technologies.

• Strong expertise in distributed systems, backend infrastructure, or high-performance platform engineering.

• Experience in optimizing inference performance across GPU workloads, focusing on latency, throughput, batching, memory utilization, and serving efficiency.

• Proficient in operating GPU infrastructure in a production setting.

• High proficiency in Python and at least one systems programming language (such as Go, Rust, or C++).

• Demonstrated ability to lead technical architecture for complex infrastructure projects.

• Exceptional communication skills with the capability to influence technical direction across engineering teams.


🏝️ Benefits

• Health insurance

People also viewed

Alight Solutions21 hours ago

Senior AI Engineer

PL flagPoland OnlyFull-timeAI Engineer
ApplyView job
Atlas Technica1 day ago

AI Engineer

UA flagUkraine OnlyFull-timeAI Engineer
ApplyView job
Creative Chaos1 day ago

Principal AI Engineer

IN flagIndia, +4 more statesFull-timeAI Engineer
ApplyView job
WCG1 day ago

Senior Director, AI Engineering

US flagUnited States OnlyFull-timeAI Engineer$151.2k – $275k/year
ApplyView job
Ford Motor Company1 day ago

AI Engineer

MX flagMexico OnlyFull-timeAI Engineer
ApplyView job
Carbon601 day ago

Principal Data, AI Architect

CA flagCanada OnlyFull-timeAI EngineerC$180k – C$220k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers