
Senior TPM – Global Reliability
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Take ownership of and enhance Slack’s incident management program from start to finish, encompassing detection, triage, response, resolution, and post-incident analysis.
• Lead the strategic evolution of incident response operations within the Customer Experience team and Salesforce’s Command Incident Center.
• Align processes, integrate tools, transfer runbooks, and facilitate cross-organizational training.
• Establish and continually refine incident severity frameworks, escalation paths, and communication protocols.
• Collaborate with Reliability leadership to assess and decrease customer-impacting incident volume and mean time to resolution.
• Oversee Slack’s reliability initiatives review, focusing on tracking, prioritizing, and implementing reliability enhancements.
• Manage load balancing and compute services programs to accommodate traffic surges and ensure smooth scaling.
• Lead capacity planning and load-shedding strategies in partnership with infrastructure engineering teams.
• Monitor and report on reliability and availability metrics, including SLOs, error budgets, and incident trends.
• Define reliability standards and SLO frameworks for agentic workloads.
• Create incident response playbooks tailored for AI failure scenarios.
• Advocate for reliability-as-a-feature in AI product development.
• Coordinate efforts across platform, infrastructure, security, and Salesforce partner teams.
• Develop scalable policies, processes, and operational rhythms.
• Create and manage timelines, risk registers, dependency maps, and executive dashboards.
• Over 8 years of experience leading technical programs in a fast-paced product or engineering organization, demonstrating progressive scope and complexity.
• Strong verbal and written communication skills.
• Adequate technical expertise to effectively converse with engineers and identify technical risks.
• Capability to work independently and communicate across various time zones.
• Exceptional organizational and interpersonal/social skills.
• Experience managing activities across multiple teams.
• Ability to analyze extensive data sets and distill them into narratives and presentations.
• Proficient in SQL for extracting large data sets into executive-level dashboards.
• Proven history of delivering complex technical projects and programs with multifunctional teams.
• At least 3 years of experience in developing and managing programs within an SRE, reliability, or infrastructure organization.
• Knowledgeable in AWS Cloud offerings or comparable cloud services.
• Direct experience with incident management programs, including severity frameworks, escalation processes, and post-incident review methodologies.
• Experience with data residency, compliance, or regulated infrastructure programs across multiple regions or jurisdictions is a significant advantage.
• Familiarity with AI/ML infrastructure, LLM serving platforms, or agentic systems is beneficial, along with demonstrated ability to rapidly acquire technical fluency in new areas.
• Experience navigating large organizational integrations is highly regarded.
• A relevant technical degree is required.
• Time off programs
• Medical insurance
• Dental insurance
• Vision insurance
• Mental health support
• Paid parental leave
• Life insurance
• Disability insurance
• 401(k)
• Employee stock purchasing program
• Reasonable accommodation during the application or recruiting process
Zayo Group
Sumsub
CertiPath
Robots & Pencils
Get handpicked remote jobs straight to your inbox weekly.