
Senior ML Systems Engineer, Inference
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Establish metrics for inference performance, including throughput, time to first token, inter-token latency, and cost per token.
• Develop tools that ensure performance measurements are precise and reproducible.
• Analyze and troubleshoot performance issues throughout the serving stack, from scheduling and memory management to kernels and interconnect.
• Enhance serving efficiency for large, advanced models in both single-node and multi-node GPU deployments.
• Transform insights into production-ready runtimes, configurations, and defaults.
• Collaborate with product and infrastructure teams to influence how inference is provided on Runpod.
• Keep an eye on the rapidly evolving inference ecosystem and assess what should be adopted, built, or contributed back.
• Identify bottlenecks in the serving engine/runtime and implement solutions when configuration tuning alone is inadequate.
• Take ownership of LLM serving performance across models, hardware generations, and workloads.
• Over 5 years of professional experience in system engineering.
• Extensive, hands-on experience with vLLM, SGLang, or a similar serving engine in production or at significant benchmark scale.
• Proficient software engineering skills in Python.
• Comfortable navigating large, performance-sensitive codebases.
• Strong comprehension of LLM inference performance, encompassing batching, memory, parallelism, and latency-throughput trade-offs.
• Experience with inference optimization methods such as quantization, speculative decoding, or distributed serving.
• Meticulous approach to benchmarking and performance analysis.
• Familiarity with GPU profiling tools.
• Ability to clearly articulate results in writing and convert them into actionable decisions.
• Must be eligible to work in the United States.
• Should not require employment visa sponsorship.
• Preferred: experience in writing or optimizing GPU kernels using CUDA or Triton.
• Preferred: contributions to inference or machine learning systems projects.
• Preferred: experience with multi-node GPU systems and high-speed networking.
• Preferred: background in a company where inference cost and latency were essential business metrics.
• Meaningful equity in a rapidly growing company; all team members receive stock options.
• Comprehensive medical, dental, and vision insurance plans.
• Flexible paid time off (PTO).
• Remote work-first policy.
• $1,200 stipend for home office setup and equipment.
• Enthusiastic team at the forefront of AI infrastructure, with a culture of learning and ownership central to the company's growth.
Quora
Amgen
Get handpicked remote jobs straight to your inbox weekly.