
Manager – Site Reliability Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Inspire the design and execution of highly available, resilient, and scalable cloud platforms utilizing AWS and cloud-native technologies
• Develop and implement reliability strategies, SLOs, SLIs, and error budgets for essential business services
• Design and manage Kubernetes-based platforms for containerized and microservices applications
• Set enterprise standards for reliability, availability, disaster recovery, resiliency testing, and operational excellence
• Collaborate with software engineering, platform engineering, security, and operations teams to enhance reliability and minimize operational risk
• Propel DevSecOps, CI/CD automation, Infrastructure as Code, and self-service platform capabilities
• Establish observability standards encompassing monitoring, logging, distributed tracing, synthetic monitoring, and operational intelligence
• Direct incident response, conduct post-incident reviews, root cause analyses, and drive continuous improvement
• Mitigate single points of failure through proactive reliability engineering and resilient architecture
• Implement automated remediation, self-healing capabilities, and proactive alerting
• Oversee capacity planning, performance engineering, availability management, and scalability assessments
• Enhance cloud resource utilization alongside FinOps and engineering teams
• Assess AIOps, Generative AI, and intelligent automation for platform reliability
• Mentor SREs, platform engineers, and software engineers
• Formulate executive-level reliability roadmaps, operational strategies, and platform investment proposals
• Counsel executive leadership on reliability strategies, operational risks, and service resiliency
• Provide technical guidance during significant incidents, outages, and high-severity escalations
• Support PCI-DSS, SOC 2, security governance, and operational resilience compliance initiatives
• Contribute to cloud transformation, platform engineering, and application modernization efforts
• Offer technical leadership and operational oversight for outsourced/offshore incident management teams
• Oversee vendor performance, incident response execution, operational effectiveness, and compliance with service-level objectives
• Bachelor’s degree in computer science, Engineering, Information Technology, or a related field
• 10+ years of experience in software engineering, cloud architecture, infrastructure engineering, or enterprise architecture
• 5+ years of practical AWS architecture and cloud transformation experience
• 5+ years of leadership or management experience
• Demonstrated success in leading large-scale cloud migrations and modernization projects
• Experience in designing and maintaining highly available, mission-critical, customer-facing platforms
• Profound knowledge of Kubernetes, container orchestration, microservices, and distributed systems
• Extensive background in DevSecOps, CI/CD pipelines, Infrastructure as Code, and automation frameworks
• Strong expertise in Site Reliability Engineering (SRE), operational excellence, and platform reliability practices
• Experience in implementing cloud governance, FinOps, and cost optimization initiatives
• Preferred: experience with large-scale enterprise, eCommerce, aviation, travel, SaaS, or high-volume digital platforms
• Preferred: experience in developing internal developer platforms and platform engineering capabilities
• AWS Professional and/or Kubernetes certifications are preferred
• Familiarity with AIOps, intelligent automation, and reliability analytics is preferred
• Knowledge of cloud networking and security, high availability and disaster recovery, performance engineering, capacity planning, observability, incident response, root cause analysis, CI/CD automation, configuration management, and reliability automation
• Medical, dental, and vision insurance coverage
• 401(k) retirement savings plans
• Paid holidays, vacation time, and sick leave
• Travel benefits on Frontier Airlines and participating partner airlines
• Buddy passes
• Discounts on travel-related expenses and select products, services, and vendors
• A hybrid work schedule for eligible roles based in Denver, Colorado
• Business casual dress code for applicable corporate and support roles
• Employee support programs and resources, including the HOPE League
Entarian
BeyondTrust
BeyondTrust
Scribe
Get handpicked remote jobs straight to your inbox weekly.