Site Reliability Engineering Lead

Posted 3 days ago

This is a fully remote position, open to applicants in Connecticut, +5 more states.

📋 Description

• Oversee, mentor, and develop a team of Site Reliability Engineers (SREs); conduct one-on-one meetings, performance evaluations, and career advancement planning.

• Take ownership of hiring, onboarding, and decisions regarding team capacity and resources.

• Establish team objectives, prioritize the backlog, and guide planning efforts.

• Promote a blameless culture in post-incident reviews and encourage collaboration across Development, Security, and Product teams.

• Spearhead reliability initiatives throughout infrastructure and services.

• Facilitate incident response efforts and drive continuous improvements in service quality.

• Advocate for automation and operational excellence within the platform.

• Assist in the creation of scalable, secure, and resilient cloud-native environments.

• Provide line management for a small to medium-sized team, including performance management, compensation, and recruitment authority.

• Conduct post-mortem reviews and ensure prompt creation of Root Cause Analyses (RCAs).

• Evaluate and resolve issues that have implications beyond the immediate team.


⛳️ Requirements

• Extensive knowledge of Kubernetes, covering cluster architecture, upgrades, autoscaling, security hardening, and troubleshooting at scale.

• Proficient in Terraform, focusing on modular Infrastructure as Code (IaC) design, state management, multi-environment provisioning, and policy-as-code.

• In-depth understanding of Azure Cloud services, including compute, networking, identity (AAD), storage, and cost optimization strategies.

• Experience in designing and scaling Continuous Integration/Continuous Deployment (CI/CD) pipelines utilizing GitHub Actions, release strategies, and rollback automation.

• Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, and managing Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.

• Strong automation capabilities aimed at reducing manual work through self-healing systems and infrastructure automation.

• Advanced skills in Python, Bash, and/or PowerShell for tooling and automation purposes.

• Comprehensive understanding of networking concepts, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking principles.

• Background in Site Reliability Engineering (SRE), DevOps, or Infrastructure roles, with experience in leading engineering teams.

• Demonstrated success in leading incident response efforts and enhancing reliability.


🏝️ Benefits

• Annual incentive bonus.

• Country-specific benefits.

• Accommodation or adjustment support during the hiring process.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers