
Staff Software Engineer, AI Inference Gateway
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Take full ownership of the AI Gateway system from start to finish, including request routing, streaming, batch/async job management, and horizontal scaling across various gateway servers.
• Develop and optimize fleet-efficiency algorithms for node autoscaling, performance evaluation, and removal of underperforming assets.
• Manage and scale inference nodes operating vLLM, llama.cpp, and similar server technologies.
• Create and uphold OpenTelemetry-based observability throughout the gateway and fleet.
• Ensure precise pay-per-token and subscription billing integrations that rely on telemetry data.
• Allocate approximately 25% of your time to benchmarking new models, quantizations, and GPU hardware.
• Collaborate with product and marketing teams to determine production priorities.
• Work together on cross-system design and integrations with Salad engineering.
• Articulate technical trade-offs to both technical and non-technical team members.
• Report directly to the CTO.
• Engage in shared after-hours support.
• Extensive production experience with Rust, particularly in high-throughput networked services.
• Practical experience managing LLM inference servers (such as vLLM, llama.cpp, TGI, TensorRT-LLM, or equivalent).
• Understanding of key-value cache behavior.
• Fundamentals of distributed systems: load balancing, backpressure, failure management, and handling unreliable nodes.
• Experience with operating systems, including metrics, tracing, and on-call responsibilities.
• Willingness to participate in on-call rotations.
• Proficient in clear written and verbal communication.
• Preferred: Familiarity with Pingora, Tokio, or proxy/gateway internals.
• Preferred: Experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm.
• Preferred: Knowledge of quantization formats such as GGUF, AWQ, GPTQ, and FP8.
• Strong expertise in at least two areas: Rust, LLM inference servers, and distributed systems.
• Unlimited paid time off (PTO).
• 75% coverage of health insurance premiums for you and your dependents.
• Dental and vision insurance included.
• 401(k) retirement plan available.
• Stock options offered.
• Company-provided computer equipment.
• $500 budget for work-from-home expenses.
• Fully remote work with flexible hours.
Terac
Echelon Services, LLC
Motive
Gametime
Get handpicked remote jobs straight to your inbox weekly.