
SRE Engineering Manager
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United Kingdom.
• Implement and enhance SRE practices throughout the team.
• Achieve quantifiable improvements in availability, performance, and incident response.
• Collaborate with Platform and LiveOps to create and operate secure, efficient, and scalable solutions.
• Equip services with metrics, tracing, and logging through observability frameworks.
• Assist in on-call readiness and conduct blameless post-mortems.
• Incorporate lessons learned from incident responses into engineering and product strategies.
• Develop and refine Infrastructure as Code (IaC) and Continuous Integration/Continuous Deployment (CI/CD) practices utilizing Terraform, GitHub Actions, Ansible, Packer, and Kubernetes.
• Establish and advance Service Level Objectives (SLOs), on-call arrangements, toil reduction strategies, and production readiness practices.
• Assess architecture proposals and prototype reliability solutions.
• Troubleshoot intricate production issues and make informed trade-off decisions regarding speed, resilience, and cost.
• Identify systemic reliability vulnerabilities and prioritize the remediation of technical debt.
• Mentor engineers while promoting psychological safety, curiosity, and accountability.
• Work alongside Security and Compliance teams to integrate SOC2, ISO27001, and PCI requirements.
• Demonstrated experience leading SRE, DevOps, or Cloud/Software engineering teams, typically comprising 5+ engineers, in a SaaS or complex multi-product setting.
• Strong technical expertise in cloud architecture (AWS, Azure, GCP, or equivalent) and hybrid or cloud migration strategies.
• Proficient in reading and reviewing application and infrastructure code.
• Comfortable with automation scripting, troubleshooting distributed system issues, and making architecture decisions based on practical experience.
• Solid foundation in observability tools and incident management methodologies, including monitoring, tracing, and alerting.
• Experience utilizing observability data to resolve actual production issues.
• Excellent communication and facilitation skills with Product, Platform, Security, Support, and LiveOps stakeholders.
• Strong mentoring capabilities.
• A passion for reliability and user experience as a cultural value.
• Experience with large-scale multi-tenant SaaS or hybrid cloud migrations (preferred).
• Familiarity with ITIL or incident command frameworks adapted for modern SRE cultures (preferred).
• Knowledge of Grafana, Datadog, ELK, and Prometheus (preferred).
• Experience working with distributed global teams (preferred).
• Background in software development and a strong understanding of the software lifecycle (preferred).
• 25 Days Annual Leave + bank holidays.
• Option to purchase up to 10 additional days.
• Days of Difference – Up to 3 extra days off for volunteering.
• Pension Contributions – 5% employer match.
• Income Protection – Up to 75% salary coverage for long-term illness.
• Life Assurance – 4x salary tax-free lump sum.
• Critical Illness Cover – £25,000 lump sum, extendable to dependents.
• Private Medical Insurance.
• Health Cash Plan – Claim back for physiotherapy, therapies, and more.
• Dental Insurance.
• Affinity Groups – Join employee-led communities.
• Bounty Bonus – Refer a friend and receive rewards.
• An inclusive and diverse workplace.
• Equal opportunity employer.
• Recruitment adjustments or accommodations available.
CDW
FAR.AI
GE Vernova
Mercury
Get handpicked remote jobs straight to your inbox weekly.