
Principal Product Manager, Augmented Memory Grid
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the AMG product roadmap, encompassing KV-cache/prefix-cache offload, memory tiering, and integrations with inference engines and orchestration layers.
• Collaborate with engineering teams to define architectural trade-offs involving GPU memory, high-performance networking, and distributed storage solutions.
• Convert inference performance constraints, such as time-to-first-token, throughput, and context length, into actionable product requirements.
• Engage with enterprise clients and GPU cloud partners to gather specifications, validate benchmarks, and prioritize features aimed at reducing cost-per-token and enhancing SLAs.
• Work alongside NVIDIA and other silicon/inference-stack partners on joint roadmap initiatives and certification processes.
• Establish and monitor TTFT, throughput, and cache-hit-rate benchmarks that illustrate AMG's value proposition.
• Assist sales and field teams with technical positioning, competitive differentiation, and support for enterprise deals.
• Practical experience in product management or engineering with LLM inference serving, including vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Ray Serve, or similar systems.
• Proficiency in KV-cache, prefix/context caching, quantization, and batching methodologies.
• Familiarity with contemporary production LLM serving practices, including context windows, multi-tenant serving, and GPU scheduling.
• Understanding of high-performance networking and distributed systems fundamentals, including RDMA, GPUDirect Storage, NVMe-oF, or equivalent technologies.
• Proven experience working directly with large enterprise clients, encompassing requirements gathering, production rollouts, and SLAs.
• Over 10 years of product management experience, preferably with a focus on infrastructure, ML platforms, or developer-oriented technical products.
• Strong leadership abilities with experience in guiding cross-functional teams.
• Capacity to influence without formal authority.
• Strategic mindset with the ability to transform goals into implementable plans.
• Exceptional communication and interpersonal skills.
• Competence in explaining intricate technical concepts to both technical and non-technical stakeholders.
• Nice-to-have: experience in a GPU cloud, inference platform, or AI infrastructure startup.
• Nice-to-have: familiarity with storage systems utilized in AI/ML workflows.
• Nice-to-have: experience collaborating with NVIDIA or other accelerator/silicon manufacturers.
• Equal opportunity employment.
• A diverse, inclusive, and authentic workplace.
• Zero tolerance for discrimination or harassment based on protected characteristics.
Allogene Therapeutics
Implus
AACN (American Association of Critical-Care Nurses)
Empower
Get handpicked remote jobs straight to your inbox weekly.