
Distinguished Engineer, Scaled Out Inferencing
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Design and implement distributed pipelines, spearheading the technical execution of high-throughput, low-latency distributed inference systems tailored for large-scale AI applications.
• Oversee the co-optimization of hardware and software, focusing on performance tuning at both the kernel and driver levels.
• Enhance GPU resource management and hardware acceleration for robust production-grade model serving.
• Direct and shape contributions to open-source projects such as Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, and Ray.
• Manage the complete model lifecycle, which includes automated deployment, versioning, and intelligent scaling across cloud and datacenter settings.
• Work collaboratively with customers, infrastructure providers, and partners to define performance and availability benchmarks.
• Steer technical planning and ongoing development throughout ideation, architecture, design, development, deployment, operations, and lifecycle management.
• Foster organizational alignment among technical teams and senior corporate leadership.
• Integrate cross-functional requirements into architecture and design while overseeing execution across various teams.
• 16+ years of experience in technical roles.
• Recent emphasis on AI infrastructure.
• Direct involvement in large-scale inference orchestration in recent positions.
• Proven experience in developing secure, highly available, and resilient production distributed systems.
• 7-10+ years of leadership experience.
• BS/MS or higher degree, or equivalent experience in systems/software engineering or related fields.
• Expertise in GPU architecture, hardware acceleration, and low-level performance tuning, including familiarity with CUDA and kernels.
• Experience with cloud-native architectures for multi-tenant model serving.
• Successful track record in delivering technically complex solutions with clear insights into resource utilization, performance, and operational metrics.
• Skill in fostering consensus and organizational alignment among technical and senior corporate leadership.
• Ability to synthesize cross-functional needs into architecture and design while guiding execution across diverse teams.
• Strong collaboration and influence capabilities.
• Practical experience in building systems that support AI/ML workloads.
• Direct experience in designing, developing, delivering, and managing secure, highly available, and scalable systems in enterprise and cloud environments.
• Proven history of creating scalable processes and extensible systems to facilitate cross-functional collaboration and operations at scale.
• Familiarity with open-source ecosystems and projects such as Dynamo, TensorRT-LLM, vLLM, SGLang, and Ray.
• Ability to collaborate and exert influence within open-source project governance.
• Competitive salaries.
• Comprehensive benefits package.
• Equity options.
• Additional benefits.
Highland Electric Fleets
Falconwood, Incorporated
Aira
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.