Senior ML Systems Engineer, Inference

atRunPodRemoteUS flagUnited StatesFull-timeMachine Learning EngineerSenior$150k – $220k/year

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Establish metrics for inference performance, including throughput, time to first token, inter-token latency, and cost per token.

• Develop tools that ensure performance measurements are precise and reproducible.

• Analyze and troubleshoot performance issues throughout the serving stack, from scheduling and memory management to kernels and interconnect.

• Enhance serving efficiency for large, advanced models in both single-node and multi-node GPU deployments.

• Transform insights into production-ready runtimes, configurations, and defaults.

• Collaborate with product and infrastructure teams to influence how inference is provided on Runpod.

• Keep an eye on the rapidly evolving inference ecosystem and assess what should be adopted, built, or contributed back.

• Identify bottlenecks in the serving engine/runtime and implement solutions when configuration tuning alone is inadequate.

• Take ownership of LLM serving performance across models, hardware generations, and workloads.


⛳️ Requirements

• Over 5 years of professional experience in system engineering.

• Extensive, hands-on experience with vLLM, SGLang, or a similar serving engine in production or at significant benchmark scale.

• Proficient software engineering skills in Python.

• Comfortable navigating large, performance-sensitive codebases.

• Strong comprehension of LLM inference performance, encompassing batching, memory, parallelism, and latency-throughput trade-offs.

• Experience with inference optimization methods such as quantization, speculative decoding, or distributed serving.

• Meticulous approach to benchmarking and performance analysis.

• Familiarity with GPU profiling tools.

• Ability to clearly articulate results in writing and convert them into actionable decisions.

• Must be eligible to work in the United States.

• Should not require employment visa sponsorship.

• Preferred: experience in writing or optimizing GPU kernels using CUDA or Triton.

• Preferred: contributions to inference or machine learning systems projects.

• Preferred: experience with multi-node GPU systems and high-speed networking.

• Preferred: background in a company where inference cost and latency were essential business metrics.


🏝️ Benefits

• Meaningful equity in a rapidly growing company; all team members receive stock options.

• Comprehensive medical, dental, and vision insurance plans.

• Flexible paid time off (PTO).

• Remote work-first policy.

• $1,200 stipend for home office setup and equipment.

• Enthusiastic team at the forefront of AI infrastructure, with a culture of learning and ownership central to the company's growth.

People also viewed

MWDN1 day ago

AI/ML Engineer

HR flagCroatia OnlyFull-timeMachine Learning Engineer
ApplyView job
Quora1 day ago

Software Engineer, New Grad – Machine Learning Platform

US flagUnited States, +1 more countryFull-timeMachine Learning Engineer$97.6k – $139k/year
ApplyView job
Amgen1 day ago

Principal Machine Learning Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer$187.4k – $253.5k/year
ApplyView job
Latitude IT Solutions | SDVOSB1 day ago

Senior AI/ML Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer$195k – $208k/year
ApplyView job
Shuru1 day ago

MLOps Engineer

IN flagIndia OnlyFull-timeMachine Learning Engineer
ApplyView job
Amgen1 day ago

Director, AI & Machine Learning

US flagUnited States OnlyFull-timeMachine Learning Engineer$242.5k – $328.1k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers