Remotery

Technical Program Manager – Provider Management

Posted 22 hours ago

This is a fully remote position, open to applicants in California.

📋 Description

• **Technical Program Manager - Provider Management** Location: Remote/SF-Hybrid · Full-Time **About Andromeda** Andromeda Cluster was established by Nat Friedman and Daniel Gross to provide early-stage startups with access to the type of scaled AI infrastructure that was previously available only to hyperscalers.

• We started with a single managed cluster, which quickly reached capacity. Since then, we have been diligently constructing the systems, network, and orchestration layer to make global AI infrastructure more accessible.

• Currently, Andromeda collaborates with top AI labs, data centers, and cloud providers to deliver computing power where it is most needed. Our platform efficiently routes training and inference jobs across global supply, enhancing flexibility and efficiency in one of the fastest-growing markets worldwide.

• Our long-term goal is to establish the liquidity layer for global AI compute. We are venturing into new areas to recruit the brightest minds in AI infrastructure, research, and engineering.

• **The Role** This position is responsible for managing our provider relationships on the execution front. You will oversee the programs that ensure a provider functions effectively within Andromeda, which includes onboarding new providers and sites, capacity rollouts, addressing hardware and fabric quality issues, and managing major incidents. You will be the point of contact for our providers regarding delivery timelines and outstanding items. In the event of a customer’s training run being interrupted due to a provider-side failure, you will act as the incident commander, coordinating the response among our SREs, the provider's engineers, and the customer until the issue is resolved.

• You will not have formal authority over provider teams; however, you will possess the plan, the escalation path, the contractual obligations, and the relationship. Mastery in utilizing all four elements is key.

• Currently, this aspect is not owned end to end; it is handled by whichever engineer has the capacity, leading to noticeable inconsistency for both providers and customers. You would be the first individual in this role, and part of your responsibility will be to define what provider program management entails at Andromeda.

• **What You’ll Do**

• - Oversee the onboarding of new providers and sites, capacity expansions, resolution of recurring hardware or network issues, and management of provider relationships.

• - Serve as the incident commander during significant provider-side incidents: assemble the appropriate personnel from our SRE team, the provider, and the affected customer, lead the response, manage communication throughout, and drive the post-incident remediation program. You will coordinate and command; our SREs will remain the technical leads making engineering decisions.

• - Maintain a comprehensive plan for each program, including milestones, responsible parties, dependencies, and risks, ensuring visibility for everyone involved, including the provider.

• - Ensure providers adhere to their commitments through structured check-ins, tracked action items, and escalation to provider leadership when commitments are not met, supported by contractual agreements.

• - Coordinate internally to prevent provider issues from causing delays. Engage SRE for validation, Engineering and Product for platform-side obstacles, and keep Sales/Customer Success updated when provider timelines impact customer commitments.

• - Identify capacity and quality risks early, before they escalate into customer-facing issues.

• - Transform your insights into playbooks and provider-facing standards, streamlining the onboarding process for subsequent sites.

• **What We’re Looking For**

• - Several years of experience managing technical programs in infrastructure. TPM or technical project management with genuine execution responsibility, preferably involving external vendors, partners, or suppliers you did not control.

• - Experience in incident management: you have led or managed production incidents involving multiple organizations, and you are comfortable with the off-hours realities that such roles entail.

• - Sufficient technical knowledge to engage with a provider's data-center engineers and our SREs on GPUs, networking, and storage. While you do not need to debug an InfiniBand fabric yourself, you should be able to follow the discussion and recognize when something is amiss.

• - A proven history of aligning teams you do not manage, including external partners, especially during challenging times when relationships may be strained.

• - Strong foundational skills: planning, risk tracking, status communication, and managing four or five programs simultaneously.

• - Calm and direct communication style. You prefer to deliver bad news promptly rather than good news late, earning the trust of both providers and internal teams.

• - Comfort with ambiguity. The processes you will follow may not yet exist; you will be responsible for creating many of them.

• **Strong Candidates May Have**

• - Direct experience working with or within neocloud, colocation, or data-center providers.

• - Background in vendor or supplier management — SLAs, commitments, escalation frameworks.

• - Familiarity with the technical stack: NVIDIA data-center GPUs, InfiniBand/RoCE, Slurm or Kubernetes.

• - Experience supporting AI research labs or other large-scale GPU customers from the consumption side.

• **What Success Looks Like** Within your first year, you will have achieved the following:

• - **Reliable provider relationships.** Our compute providers trust you and communicate transparently with you, so when they provide updates on capacity availability, we can depend on that information for planning. Capacity forecasting is reliable rather than based on hope.

• - **Efficient onboarding processes.** New providers and sites are integrated on a predictable timeline with minimal surprises. Each onboarding experience is improved as you systematize the process.

• - **A partner playbook that enhances network scalability.** You will have developed standards and playbooks that potential providers can use to align their infrastructure with our requirements. This will transform the qualifying and onboarding of new providers into a repeatable process rather than a unique project with each instance.

• - **Trust built internally across teams you do not manage.** SRE, Engineering, Product, and Sales/CS will rely on your insights regarding provider health and timelines, and they will count on you to proactively identify risks and coordinate responses when necessary.

• **Why You’ll Love It Here**

• - **High-growth environment:** Join a company at the forefront of the AI infrastructure surge by getting in early.

• - **Ownership:** As the first TPM hire for the solutions engineering team, you will have the opportunity to establish this function from the ground up.

• - **Competitive compensation:** Plus meaningful equity.

• - **Comprehensive benefits:** For you and your dependents, including healthcare, dental, and vision coverage, a 401(k), and unlimited PTO.


⛳️ Requirements

• Several years of experience managing technical programs in infrastructure.

• Experience in incident management.

• Adequate technical knowledge of GPUs, networking, and storage.

• Proven track record of aligning teams and external partners.

• Strong foundational skills in planning, risk tracking, and communication.

• Calm, direct communication style.

• Comfort with ambiguity.


🏝️ Benefits

• Competitive compensation.

• Meaningful equity.

• Comprehensive benefits for you and your dependents, including healthcare, dental, and vision coverage.

• 401(k).

• Unlimited PTO.

People also viewed

Virtasant2 hours ago

Technical Program Manager – Capacity Management

CA flagCanada OnlyFull-timeTechnical Program Manager
ApplyView job
comrce group2 hours ago

Technical Program Manager – m/f/d

DE flagGermany OnlyFull-timeTechnical Program Manager€50k – €60k/year
ApplyView job
Acryl Data2 hours ago

Technical Program Manager – Release Management

US flagUnited States OnlyFull-timeTechnical Program Manager
ApplyView job
Cisco2 hours ago

Technical Program Manager

US flagMassachusetts, +1 more stateFull-timeTechnical Program Manager$135.4k – $171.4k/year
ApplyView job
Google Fiber2 hours ago

Lead Network Technical Program Manager

US flagUnited States OnlyFull-timeTechnical Program Manager$164k – $241.1k/year
ApplyView job
Akamai Technologies22 hours ago

Senior Technical Program Manager

US flagMassachusetts OnlyFull-timeTechnical Program Manager$119.6k – $215.4k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers