Manager – Site Reliability Engineer

Posted 6 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Inspire the design and execution of highly available, resilient, and scalable cloud platforms utilizing AWS and cloud-native technologies

• Develop and implement reliability strategies, SLOs, SLIs, and error budgets for essential business services

• Design and manage Kubernetes-based platforms for containerized and microservices applications

• Set enterprise standards for reliability, availability, disaster recovery, resiliency testing, and operational excellence

• Collaborate with software engineering, platform engineering, security, and operations teams to enhance reliability and minimize operational risk

• Propel DevSecOps, CI/CD automation, Infrastructure as Code, and self-service platform capabilities

• Establish observability standards encompassing monitoring, logging, distributed tracing, synthetic monitoring, and operational intelligence

• Direct incident response, conduct post-incident reviews, root cause analyses, and drive continuous improvement

• Mitigate single points of failure through proactive reliability engineering and resilient architecture

• Implement automated remediation, self-healing capabilities, and proactive alerting

• Oversee capacity planning, performance engineering, availability management, and scalability assessments

• Enhance cloud resource utilization alongside FinOps and engineering teams

• Assess AIOps, Generative AI, and intelligent automation for platform reliability

• Mentor SREs, platform engineers, and software engineers

• Formulate executive-level reliability roadmaps, operational strategies, and platform investment proposals

• Counsel executive leadership on reliability strategies, operational risks, and service resiliency

• Provide technical guidance during significant incidents, outages, and high-severity escalations

• Support PCI-DSS, SOC 2, security governance, and operational resilience compliance initiatives

• Contribute to cloud transformation, platform engineering, and application modernization efforts

• Offer technical leadership and operational oversight for outsourced/offshore incident management teams

• Oversee vendor performance, incident response execution, operational effectiveness, and compliance with service-level objectives


⛳️ Requirements

• Bachelor’s degree in computer science, Engineering, Information Technology, or a related field

• 10+ years of experience in software engineering, cloud architecture, infrastructure engineering, or enterprise architecture

• 5+ years of practical AWS architecture and cloud transformation experience

• 5+ years of leadership or management experience

• Demonstrated success in leading large-scale cloud migrations and modernization projects

• Experience in designing and maintaining highly available, mission-critical, customer-facing platforms

• Profound knowledge of Kubernetes, container orchestration, microservices, and distributed systems

• Extensive background in DevSecOps, CI/CD pipelines, Infrastructure as Code, and automation frameworks

• Strong expertise in Site Reliability Engineering (SRE), operational excellence, and platform reliability practices

• Experience in implementing cloud governance, FinOps, and cost optimization initiatives

• Preferred: experience with large-scale enterprise, eCommerce, aviation, travel, SaaS, or high-volume digital platforms

• Preferred: experience in developing internal developer platforms and platform engineering capabilities

• AWS Professional and/or Kubernetes certifications are preferred

• Familiarity with AIOps, intelligent automation, and reliability analytics is preferred

• Knowledge of cloud networking and security, high availability and disaster recovery, performance engineering, capacity planning, observability, incident response, root cause analysis, CI/CD automation, configuration management, and reliability automation


🏝️ Benefits

• Medical, dental, and vision insurance coverage

• 401(k) retirement savings plans

• Paid holidays, vacation time, and sick leave

• Travel benefits on Frontier Airlines and participating partner airlines

• Buddy passes

• Discounts on travel-related expenses and select products, services, and vendors

• A hybrid work schedule for eligible roles based in Denver, Colorado

• Business casual dress code for applicable corporate and support roles

• Employee support programs and resources, including the HOPE League

People also viewed

Entarian1 day ago

DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $200k/year
ApplyView job
BeyondTrust1 day ago

Senior Manager, Site Reliability Engineer – FedRAMP, AWS GovCloud

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BeyondTrust1 day ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Scribe1 day ago

Senior DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $240k/year
ApplyView job
Hypertegrity AG1 day ago

Senior DevOps Engineer – Smart City Open Source

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CACI International Inc1 day ago

Senior DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$98.5k – $206.8k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers