
Inference Infrastructure Architect
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in China.
• Efficiently operate and expand Telnyx’s B300 GPU fleet.
• Optimize inference throughput per GPU-dollar while adhering to latency and reliability service level objectives (SLOs).
• Continuously strive to reduce the cost per token.
• Construct serverless serving pools utilizing vLLM/SGLang, continuous batching, prefix caching, low-precision serving, and Mixture of Experts (MoE) parallelism.
• Develop the fleet layer with llm-d/NVIDIA Dynamo on Kubernetes Gateway API, implementing KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving.
• Establish Kubernetes-on-bare-metal infrastructure featuring GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers, and OpenStack Ironic.
• Implement model distribution, warm pools, and inference-metric-driven autoscaling.
• Create dedicated tenant pools, GPU-hour metering, latency SLOs, private networking, adapter versioning, canary rollouts, and rollback processes.
• Integrate observability using DCGM, Prometheus, and OpenTelemetry.
• Conduct capacity planning through roofline analysis, batching curves, utilization, and cost-per-token metrics.
• Select, benchmark, and integrate production components.
• Contribute upstream to the open-source serving stack.
• Provide a brief application example outlining an improved inference system, identifying bottlenecks, interventions, and measured results.
• Experience in production LLM serving under significant traffic and latency constraints.
• Proficient in operating Kubernetes on GPU fleets from start to finish, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, GitOps rollouts, and pod-placement debugging.
• Strong operational command of vLLM or SGLang, encompassing deployment, tuning, upgrades, tensor parallelism, execution parallelism, quantization, batching, KV-cache settings, prefix caching, and disaggregation.
• System-level performance engineering experience using engine and DCGM metrics, roofline analysis, and numerical deployment sizing.
• Proficient in Python and Go for automation tasks.
• Solid understanding of Linux, networking, and storage.
• Capability to identify CUDA-related issues and escalate them appropriately.
• Ability to engage with the community in both English and Chinese.
• Competence in writing effective runbooks and designing documents.
• Experience with inference platforms in large-scale organizations is preferred.
• Contributions to or extensive use of vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy are advantageous.
• Participation in the CNCF/OpenInfra community is valued.
• Experience in fine-tuning/RL infrastructure, real-time voice latency work, and multi-region/data residency is a plus.
• Opportunity to manage a globally expanding B300 fleet from start to finish.
• Ownership of a greenfield inference platform and the chance to assist in building the China team.
• An open-source-first work environment.
• Support for conference travel.
• Remote work with a global, asynchronous-friendly team.
• No relocation necessary.
• Visa sponsorship available for relocation to hiring entities in the Netherlands, United States, Ireland, or Saudi Arabia.
Conduent
HavocAI
Software Mind
Cint
Get handpicked remote jobs straight to your inbox weekly.