
Lead Engineer – Site Reliability
Posted Sep 16

Posted Sep 16
This is a fully remote position, open to applicants in United States.
• Oversee the design and execution of highly available, resilient, and scalable AWS cloud infrastructures.
• Develop and implement reliability engineering strategies, including SLOs, SLIs, and error budgets.
• Architect and manage Kubernetes-based platforms that support containerized and microservices applications.
• Set enterprise standards for reliability, availability, disaster recovery, resiliency testing, and operational excellence.
• Collaborate with software engineering, platform engineering, security, and operations teams to enhance reliability and minimize operational risks.
• Promote DevSecOps, CI/CD automation, Infrastructure as Code, and the adoption of self-service platforms.
• Establish observability standards that encompass monitoring, logging, distributed tracing, synthetic monitoring, and operational intelligence.
• Lead incident response efforts, conduct post-incident reviews, root cause analyses, and drive continuous improvement.
• Proactively eliminate single points of failure through reliability engineering and resilient architecture.
• Design automated remediation processes, self-healing capabilities, and proactive alert systems.
• Manage capacity planning, performance engineering, availability oversight, and scalability evaluations.
• Enhance cloud resource utilization alongside FinOps and engineering teams.
• Assess AIOps, Generative AI, and intelligent automation solutions.
• Mentor SREs, platform engineers, and software engineers.
• Create executive-level reliability roadmaps, operational strategies, and platform investment proposals.
• Counsel executive leadership on reliability strategies, operational risks, and service resiliency.
• Provide technical guidance during major incidents and high-severity escalations.
• Support compliance initiatives for PCI-DSS, SOC 2, security governance, and operational resilience.
• Contribute to cloud transformation, platform engineering, and application modernization projects.
• Advocate for reliability, automation, operational ownership, and continuous improvement.
• A Bachelor’s degree in computer science, engineering, information technology, or a related field.
• More than 10 years of experience in software engineering, cloud architecture, infrastructure engineering, or enterprise architecture.
• At least 5 years of hands-on experience in AWS architecture and cloud transformation.
• Demonstrated success in leading large-scale cloud migrations and modernization efforts.
• Experience in designing and maintaining highly available, mission-critical, customer-facing platforms.
• In-depth knowledge of Kubernetes, container orchestration, microservices, and distributed systems.
• Extensive experience with DevSecOps, CI/CD pipelines, Infrastructure as Code, and automation frameworks.
• Strong understanding of Site Reliability Engineering, operational excellence, and platform reliability practices.
• Experience in implementing cloud governance, FinOps, and cost optimization strategies.
• Background in supporting large-scale enterprise platforms, including eCommerce, aviation, travel, SaaS, or high-volume digital services.
• Experience in building internal developer platforms and enhancing platform engineering capabilities.
• AWS Professional and/or Kubernetes certifications.
• Familiarity with AIOps, intelligent automation, and reliability analytics.
• Knowledge of AWS, Kubernetes/EKS, Terraform, CloudFormation, cloud networking and security, high availability, disaster recovery, performance engineering, capacity planning, SRE, SLIs, SLOs, error budgets, chaos engineering, incident response, observability, monitoring, alerting, distributed tracing, centralized logging, synthetic monitoring, DevSecOps, CI/CD, configuration management, and reliability automation.
• Medical, dental, and vision insurance coverage.
• 401(k) retirement savings plans.
• Paid holidays, vacation time, and sick leave.
• Travel benefits with Frontier Airlines and participating partner airlines.
• Eligibility for buddy passes according to program rules.
• Discounts on travel-related services and select products, services, and vendors.
• A hybrid work schedule for eligible headquarters roles located in Denver, Colorado.
• Business casual dress options for appropriate corporate and support positions.
• Employee support programs and resources, including the HOPE League.
• Remote work environment.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.