
Site Reliability Engineering Technical Leader
Posted Aug 24

Posted Aug 24
This is a fully remote position, open to applicants in Arizona, +4 more states.
• Establish the technical standards for judgment, architectural insight, and operational excellence throughout the organization
• Take ownership of the technical strategy for addressing, learning from, and mitigating critical incidents
• Act as the primary escalation point during P1/P2 events
• Make high-stakes, risk-informed decisions to safeguard customer uptime and trust
• Manage the most intricate customer environments
• Influence automation strategies and infrastructure architecture choices with regional and enterprise-wide implications
• Impact the development of Splunk Cloud operational procedures
• Collaborate with Engineering, Customer Success, Release Management, Account Managers, and various cross-functional teams
• Offer expert guidance on complex configurations, upgrades, feature rollouts, migrations, and operational management
• Convert customer feedback into quantifiable enhancements to the cloud service experience
• Develop frameworks that prevent recurrence and address operational process and runbook deficiencies
• Enhance the technical capabilities of senior engineers and cultivate senior technical talent without formal management authority
• Bachelor’s degree with 12 years of relevant experience, Master’s degree with 8 years, or PhD with 5 years
• Over 7 years of experience in SRE, cloud operations, and systems engineering with Linux administration
• More than 6 years of hands-on experience with AWS, GCP, or Azure
• At least 5 years of experience in a scripting or automation language such as Python or Go/Golang
• 5+ years of experience in communication across engineering, customer-facing, and senior leadership demographics
• 4+ years of experience leading post-mortems and conducting root cause analyses for high-severity incidents
• Proven experience in driving systemic enhancements and resolving process gaps
• History of managing high-severity customer escalations and critical production challenges across AWS, GCP, and Azure
• Enterprise-level experience encompassing architecture, deployment, infrastructure modifications, migrations, and operational management
• Proficiency in Linux systems administration and large-scale distributed systems architecture
• Experience in operating and refining monitoring, alerting, and observability systems
• Outstanding communication and technical leadership abilities
• Previous experience in customer-facing technical advisory roles such as Professional Services, Global Services, or Technical Account Management
• Medical, dental, and vision insurance
• 401(k) plan with a Cisco matching contribution
• Paid parental leave
• Short- and long-term disability coverage
• Basic life insurance
• Cisco restricted stock unit grants may be available, vesting after continued employment
• 10 paid holidays each full calendar year
• 1 floating holiday for non-exempt employees
• Paid day off for employee birthdays
• Paid year-end holiday shutdown
• 4 paid personal wellness days
• 16 days of paid vacation per full calendar year for non-exempt employees
• Flexible vacation time off with no defined limit for eligible exempt employees
• 80 hours of sick time off provided on hire date and each January 1st thereafter
• Up to 80 hours of unused sick time carried forward annually
• Additional paid time off for critical or emergency family matters
• Optional 10 paid volunteer days each full calendar year
• Annual bonuses for non-sales positions, subject to Cisco policies
• Performance-based incentive pay for employees on sales plans
FCamara Consulting & Training
Sequoia Connect
FundCount
GE Vernova
Get handpicked remote jobs straight to your inbox weekly.