Site Reliability Engineering Lead

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead and mentor a team of Site Reliability Engineers.

• Define and drive the SRE strategy, standards, best practices, and operational frameworks.

• Enhance application reliability, scalability, security, performance, and resilience across product and platform teams.

• Establish and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.

• Oversee major incident management, root cause analysis, problem management, and post-incident reviews.

• Propel cloud modernization and facilitate application migrations to Azure, AWS, and containerized environments.

• Advocate for automation and Infrastructure as Code practices utilizing Terraform, GitHub, GitLab, Jenkins, and Ansible.

• Develop and implement observability strategies encompassing metrics, logs, traces, alerting, and dashboards.

• Collaborate with security and compliance teams to meet enterprise requirements.

• Lead architecture reviews and steer cloud-native, resilient application designs.

• Manage capacity planning, performance optimization, cost management, and operational efficiency.

• Establish engineering guardrails, governance controls, and standards for production deployment.

• Support the organizational transformation towards DevOps and SRE practices.

• Manage operational risk and ensure preparedness for business continuity and disaster recovery.

• Prioritize reliability enhancements and platform investments in collaboration with stakeholders.

• Foster relationships with product owners, engineering leaders, vendors, and business partners.

• Conduct resource planning while assisting with hiring, onboarding, and career development.

• Set team objectives aligned with business and technology strategies.

• Promote accountability, innovation, and operational excellence within the team.

• Serve as a subject matter expert and trusted advisor in reliability engineering.

• Facilitate cross-team collaboration on strategic initiatives.


⛳️ Requirements

• 8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering.

• 2+ years of leadership or people management experience in directing engineering teams.

• Bachelor's degree in Computer Science, Engineering, Information Systems, or an equivalent practical experience.

• Strong leadership experience in managing technical engineering teams.

• Profound expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering.

• Extensive experience with Azure and/or AWS cloud platforms.

• Strong understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud-native architectures.

• Proficiency with Infrastructure as Code tools such as Terraform and Ansible.

• Solid background in observability platforms like Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies.

• Experience managing large-scale production environments with stringent availability requirements.

• Thorough understanding of security, compliance, networking, and cloud governance principles.

• Experience designing highly available, fault-tolerant, and resilient systems.

• Strong proficiency in at least one scripting or programming language such as Python, Go, PowerShell, Bash, or C#.

• Familiarity with CI/CD pipelines and software delivery automation.

• Exceptional troubleshooting and problem-solving abilities.

• Excellent communication and stakeholder management skills.

• Strong documentation and presentation capabilities.

• Ability to influence technical direction across multiple engineering organizations.

• Azure, AWS, Kubernetes, Terraform, or related certifications are preferred.

• Proven track record of leading reliability and operational excellence initiatives in large-scale enterprise environments.


🏝️ Benefits

• Annual incentive bonus eligibility.

• Country-specific benefits.

• Reasonable accommodations and adjustments during the hiring process.

• Equal opportunity employment.

People also viewed

VALCE Talent Solutions14 hours ago

AWS DevOps

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
4Pharma Ltd16 hours ago

Senior Dev Ops Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
MTP Brasil16 hours ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BlackSky16 hours ago

Principal DevSecOps Architect

US flagVirginia, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)$190k – $215k/year
ApplyView job
BPCS, Comprehensive marketing solutions, ltd.16 hours ago

DevSecOps Engineer

US flagWashington OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $105k/year
ApplyView job
FTI - Frontier Technology Inc.19 hours ago

Senior DevOps Engineer

US flagAlabama, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers