
Senior Technical Program Manager, AI Infrastructure – Capacity Operations
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California.
• Take ownership of the intake, triage, routing, and quality standards for accelerator-capacity requests across engineering and research teams.
• Conduct quarterly and annual demand forecasts, manage change control, prepare for reviews, and ensure follow-through.
• Develop capacity-planning and allocation reviews based on demand, supply, commitments, readiness, workload timing, and business priorities.
• Monitor capacity from the point of request and forecast through delivery, readiness, assignment, and effective utilization.
• Coordinate infrastructure dependencies such as access, storage, data movement, networking, readiness checks, migration timing, and provisioning tickets.
• Maintain dashboards and ensure source-data quality, which includes conducting freshness checks, reconciliation, ownership of missing inputs, and retiring repetitive manual reporting.
• Oversee operating cadences, agendas, action logs, tracking dependencies, maintaining decision records, managing risk registers, and resolving blocking issues.
• Prepare succinct leadership reports that cover facts, risks, decisions, options, recommended actions, owners, and due dates.
• Lead initiatives in tooling, automation, and decision-support with human approval gates, audit evidence, and secure operating controls.
• Influence collaboration across research, platform engineering, infrastructure, data, finance, operations, and leadership teams.
• Bachelor’s, Master’s, or PhD in Electrical Engineering, Computer Science, Computer Engineering, or a related field, or equivalent experience.
• Over 7 years of technical program management or closely related experience in AI/ML platforms, distributed systems, cloud infrastructure, compute capacity, or demanding engineering environments.
• Proven experience managing recurring operational programs that involve scarce-resource trade-offs, multiple cadences, and executive clarity during overlapping peak periods.
• Strong command of program mechanics including intake, forecasting, review preparation, dependency management, documentation, risk records, action closure, and status communication.
• Sufficient technical proficiency to comprehend infrastructure constraints, review requirements and metrics, and collaborate effectively with engineers and researchers.
• Data proficiency to assess source quality, interpret dashboards, reconcile differing views, define operational measures, and pinpoint missing ownership.
• A history of transforming ambiguous cross-functional efforts into robust operating systems with defined owners, decisions, and critical issue pathways.
• Exceptional written and verbal communication skills.
• Capability to influence without authority among collaborators in research, platform engineering, infrastructure, data, finance, operations, and leadership.
• Experience in GPU or accelerator-capacity planning, utilization programs, cluster operations, workload bring-up, or large-scale AI training and inference environments.
• Familiarity with demand and supply roadmaps, normalization across various accelerator types, allocation reviews, quota or priority management, and capacity migrations.
• Understanding of batch scheduling, cluster management, and environments related to cloud or data centers.
• Experience with observability or dashboard platforms, data-quality controls, and automated infrastructure or service-operations reporting.
• Proven experience in driving API-enabled workflow automation or decision-support tools with human approval gates, audit evidence, and secure operating controls.
• Equity
• Benefits
Alteryx
Evergrid.ai
Prove
Symbotic
Get handpicked remote jobs straight to your inbox weekly.