Remotery

Engineering Manager – Site Reliability

Posted Aug 4

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Develop the framework that enables teams to establish SLOs and error budgets, providing education and support to product teams.

• Take ownership of blameless post-mortems and implement root-cause solutions.

• Conduct production readiness reviews for new releases.

• Enhance Datadog observability coverage, which includes alerts, dashboards, and on-call pages.

• Recruit and build the Site Reliability team from scratch.

• Establish the on-call rotation and incident management procedures.

• Define SLOs and create an incident review process.

• Integrate reliability practices into the software development lifecycle to minimize recurring incidents.


⛳️ Requirements

• 7 to 10 years of engineering experience, with a minimum of 3 years in direct management of engineers.

• Proven history of hiring and developing engineers, including leveling up or promoting team members.

• Hands-on experience in production or reliability engineering.

• Previous experience carrying a pager; this is not an entry-level management position.

• Strong expertise in AWS and Terraform.

• Comfortable working within a managed infrastructure-as-code pipeline.

• Experience in building or managing an on-call rotation and incident management processes.

• Strong background in platform observability.

• Excellent communication skills for a technical, cross-team audience.

• Proven track record of integrating reliability practices into product engineering teams.

• Proficient with AI-assisted development tools such as Claude Code and Cursor.

• Capable of rigorously reviewing AI-generated pull requests.

• Familiarity with PCI DSS, SOC 2, and ISO 27001 environments relevant to the stack/context.


🏝️ Benefits

• Competitive salary package.

• Choice of preferred operating system: Windows or Mac.

• Flat organizational structure and open communication in a relaxed yet professional environment.

• Opportunity to nurture your skills within a dynamic team with ambitious objectives.

• Flexibility to work remotely.

• Pliant Card with a monthly allowance to explore the product and enjoy meals with colleagues.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers