
Vice President, Global Production Operations – Reliability
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Define and implement Everbridge's worldwide production operations and reliability strategy.
• Lead and cultivate global Site Reliability Engineering (SRE) and DevOps Reliability Engineering (DRE) teams tasked with ensuring platform reliability and operational engineering.
• Take ownership of operational excellence for Everbridge's cloud platform, which is based on AWS and Kubernetes, ensuring it is scalable, resilient, secure, and high-performing.
• Establish service reliability standards, including Service Level Objectives (SLOs), error budgets, production readiness, capacity planning, observability, and operational automation.
• Oversee incident management, change governance, release management, and conduct post-incident reviews.
• Lead initiatives for disaster recovery planning, resilience testing, business continuity, and operational readiness within global production environments.
• Collaborate with Engineering, Product, Security, Customer Support, and Customer Success teams to integrate reliability into the software development lifecycle and enhance customer outcomes.
• Advocate for automation, cloud-native engineering practices, and a culture of continuous improvement.
• Provide executive leadership and reporting on operational health, reliability metrics, customer-impact incidents, and strategic initiatives.
• 15+ years of experience in production operations, cloud infrastructure, platform engineering, Site Reliability Engineering (SRE), or other related technology leadership roles.
• Proven track record in leading global production operations for large-scale, mission-critical SaaS or cloud platforms.
• Extensive knowledge of AWS, Kubernetes, cloud-native architectures, distributed systems, and high-availability environments.
• Strong grasp of SRE principles, incident management, observability, disaster recovery, change management, CI/CD, and Infrastructure as Code.
• Demonstrated achievement in building and managing high-performing global engineering and operations teams.
• Exceptional executive communication skills, stakeholder management, and cross-functional leadership abilities.
• Preferred: experience in mission-critical sectors such as public safety, critical communications, healthcare, financial services, security, or enterprise SaaS.
• Preferred: familiarity with contemporary SRE practices, progressive delivery, platform engineering, service mesh, and cloud cost optimization.
• Preferred: understanding of ISO 27001, SOC 2, NIST, or FedRAMP frameworks.
• Healthcare benefits
• Dental benefits
• Parental planning benefits
• Mental health benefits
• Disability income benefits
• Life and AD&D insurance
• 401(k) plan and matching
• Paid time off
• Fitness reimbursements
• Variable compensation may also be included
Ellit Groups
RethinkFirst
Blue Acorn iCi
Get handpicked remote jobs straight to your inbox weekly.