
Site Reliability Engineer
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in United Kingdom.
• Actively monitor and analyze the performance of the platform.
• Partner with engineering teams to identify and resolve performance bottlenecks while ensuring scalability.
• Aid engineering teams in the implementation and assessment of Service Level Objectives (SLOs).
• Enhance observability through monitoring, alerting, and dashboard creation using tools like DataDog or Prometheus.
• Collaborate with various teams to guarantee effective monitoring and comprehensive coverage.
• Ensure that services maintain high availability and resilience.
• Advocate for best practices in high-availability design.
• Develop runbooks and conduct game sessions to validate disaster recovery plans, high availability, and backup strategies.
• Evaluate capacity and plan for scaling to meet current and future business demands.
• Strategize and execute scalable solutions in partnership with the Head of Platform Engineering and Head of Site Reliability Engineering (SRE).
• Work with the Platform team, feature teams, second-line support, and stakeholders to deliver excellent customer service and integrate SRE practices.
• Address and resolve incidents promptly, ensuring rapid recovery and minimizing downtime.
• Engage in blameless postmortems to determine root causes and corrective measures.
• Create and maintain playbooks and documentation.
• Experience in performance monitoring and analysis.
• Proficiency in capacity planning.
• Skills in scripting and automation with relevant technologies.
• Experience with Infrastructure as Code, particularly using Terraform.
• Knowledge of relational database technologies and cloud versions, such as AWS Aurora.
• Familiarity with messaging and distributed asynchronous workloads.
• Experience with nginx or comparable technologies.
• Understanding of SRE processes.
• Awareness of DevOps principles, including the three ways and five ideals.
• Bonus: Experience with other database technologies and cloud platforms.
• Bonus: Previous experience with enterprise solutions operating at scale.
• Bonus: Familiarity with Kanban and Agile development methodologies.
• Bonus: Experience with containerization technologies like Docker.
• Bonus: Knowledge of software best practices, including Refactoring, Clean Code, Domain-Driven Design, and Test-Driven Development.
• A dedicated wellbeing team promoting mindfulness, lunch and learns, manager training, mental health first aid training, and more.
• 32 days of holiday plus Bank Holidays (25 days of annual leave plus 7 additional company-wide days).
• Life Assurance coverage equivalent to 3 times the annual salary.
• Comprehensive wellness benefits through AIG Smart Health, including a 24/7 virtual GP service, mental health support, counseling, and personalized health checks.
• Private Dental Insurance through Bupa.
• Salary sacrifice Pension plan provided by Scottish Widows.
• Enhanced maternity and adoption leave (20 weeks of full pay) and paternity leave (6 weeks of full pay).
• Five complimentary return-to-work maternity coaching sessions.
• Access to Calm and Bippit for financial wellbeing coaching.
• Flexible working arrangements.
• Social committees organizing team, office, and company-wide events.
• Opportunity to collaborate with a passionate team and witness the impact of your contributions.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.