Inference Infrastructure Architect

Posted 1 day ago

This is a fully remote position, open to applicants in China.

📋 Description

• Efficiently operate and expand Telnyx’s B300 GPU fleet.

• Optimize inference throughput per GPU-dollar while adhering to latency and reliability service level objectives (SLOs).

• Continuously strive to reduce the cost per token.

• Construct serverless serving pools utilizing vLLM/SGLang, continuous batching, prefix caching, low-precision serving, and Mixture of Experts (MoE) parallelism.

• Develop the fleet layer with llm-d/NVIDIA Dynamo on Kubernetes Gateway API, implementing KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving.

• Establish Kubernetes-on-bare-metal infrastructure featuring GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers, and OpenStack Ironic.

• Implement model distribution, warm pools, and inference-metric-driven autoscaling.

• Create dedicated tenant pools, GPU-hour metering, latency SLOs, private networking, adapter versioning, canary rollouts, and rollback processes.

• Integrate observability using DCGM, Prometheus, and OpenTelemetry.

• Conduct capacity planning through roofline analysis, batching curves, utilization, and cost-per-token metrics.

• Select, benchmark, and integrate production components.

• Contribute upstream to the open-source serving stack.

• Provide a brief application example outlining an improved inference system, identifying bottlenecks, interventions, and measured results.


⛳️ Requirements

• Experience in production LLM serving under significant traffic and latency constraints.

• Proficient in operating Kubernetes on GPU fleets from start to finish, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, GitOps rollouts, and pod-placement debugging.

• Strong operational command of vLLM or SGLang, encompassing deployment, tuning, upgrades, tensor parallelism, execution parallelism, quantization, batching, KV-cache settings, prefix caching, and disaggregation.

• System-level performance engineering experience using engine and DCGM metrics, roofline analysis, and numerical deployment sizing.

• Proficient in Python and Go for automation tasks.

• Solid understanding of Linux, networking, and storage.

• Capability to identify CUDA-related issues and escalate them appropriately.

• Ability to engage with the community in both English and Chinese.

• Competence in writing effective runbooks and designing documents.

• Experience with inference platforms in large-scale organizations is preferred.

• Contributions to or extensive use of vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy are advantageous.

• Participation in the CNCF/OpenInfra community is valued.

• Experience in fine-tuning/RL infrastructure, real-time voice latency work, and multi-region/data residency is a plus.


🏝️ Benefits

• Opportunity to manage a globally expanding B300 fleet from start to finish.

• Ownership of a greenfield inference platform and the chance to assist in building the China team.

• An open-source-first work environment.

• Support for conference travel.

• Remote work with a global, asynchronous-friendly team.

• No relocation necessary.

• Visa sponsorship available for relocation to hiring entities in the Netherlands, United States, Ireland, or Saudi Arabia.

People also viewed

Conduent1 day ago

Windows Infrastructure Engineer

US flagMaryland OnlyFull-timeInfrastructure Engineer$85.5k – $111k/year
ApplyView job
HavocAI1 day ago

Data and ML Infrastructure Engineer

US flagRhode Island OnlyFull-timeInfrastructure Engineer$150k – $185k/year
ApplyView job
Software Mind1 day ago

Senior Python Developer, Infrastructure

RO flagRomania, +1 more countryFull-timeInfrastructure Engineer
ApplyView job
Cint1 day ago

Senior Cloud Infrastructure Engineer

ES flagSpain, +1 more countryFull-timeInfrastructure Engineer
ApplyView job
Cint1 day ago

Senior Cloud Infrastructure Engineer

ES flagSpain OnlyFull-timeInfrastructure Engineer
ApplyView job
Atmosera1 day ago

Cloud Infrastructure Engineer

AR flagArgentina, +2 more countriesFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers