
Principal Product Manager, AI Infrastructure, Orchestration
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in United States, +4 more locations.
• Take charge of the deployment and workload API, which includes managing the resource model, lifecycle semantics, versioning, backward compatibility, and customer-facing error responses.
• Establish placement and capacity across nodes and accelerators, addressing quota, priority, and contention management across multiple tenants.
• Manage autoscaling signals, cold-start and scale-to-zero economics, headroom policies, and cost-versus-latency considerations.
• Oversee agent runtime decisions related to execution location, duration, isolation, tool calls, and state persistence during restarts or evictions.
• Define ingress, routing, request-aware load balancing, tenancy boundaries, and private connectivity for both model and agent endpoints.
• Manage governance and auditing features concerning deployments, invocations, policies, access control, and audit durability.
• Develop metering, quota, tenant attribution, inference packaging, and pricing strategies.
• Ensure reliability SLOs, error budgets, and operational diagnostics for degraded deployments are met.
• Maintain the roadmap for the workload and agent runtime layer across two engineering teams.
• Collaborate with product teams developing on the platform.
• Conduct design reviews, establish API contracts, and utilize production data effectively.
• A minimum of 6 years of experience in product management focused on infrastructure, developer platforms, or cloud services.
• At least 3 years of experience with Kubernetes-based or distributed systems products.
• Candidates for principal roles should possess over 9 years of experience and have managed a platform layer utilized by other product teams.
• In-depth technical knowledge of GPU and accelerator behavior, encompassing topology-aware placement, fractional and time-sliced sharing, MIG, device plugins, driver and container-runtime integration, memory limitations, and utilization economics.
• Extensive understanding of Kubernetes, including the API server and scheduler, controllers and CRDs, operators, admission and RBAC, device plugins, resource requests and limits, node pools, and pod scheduling failures.
• Experience with multi-tenancy, including isolation models, noisy neighbor issues, quota and fairness management, and security-review-ready tenancy designs.
• Proven experience managing a public or platform API.
• Strong technical writing and prototyping skills, including the ability to produce documentation, in-depth analyses, public postings, API references, or functional prototypes.
• Ability to work effectively with matrixed engineering teams without direct reports.
• A BS or MS in Computer Science or a related technical discipline, or equivalent practical experience as a software, platform, or infrastructure engineer.
• Preferred qualifications include expertise in service networking, modern serving stacks like vLLM, long-running and agentic workload patterns, customer-managed/air-gapped/sovereign deployments, regulated industry experience, and contributions to CNCF or open-source projects.
• Capability to independently develop prototypes, evaluation harnesses, agents connected to actual services, or other functional tools.
• Medical, Dental & Vision Insurance
• Flexible Time Off Program
• Paid Holidays
• Paid Parental Leave
• Global Employee Assistance Program (EAP)
• Competitive pay
• Reasonable accommodations for applicants with physical and mental disabilities
Vanilla
Duck Creek Technologies
Sumsub
Get handpicked remote jobs straight to your inbox weekly.