
Senior Site Reliability Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in United Kingdom, +3 more countries.
• Enhance production reliability and strengthen system resilience within a team focused on Site Reliability Engineering (SRE).
• Advocate for high-quality work and adherence to industry best practices.
• Engage in communication with teams and stakeholders throughout all phases.
• Introduce innovative ideas and motivate others.
• Tackle intricate technical challenges.
• Operate across various technologies in a rapidly evolving sector.
• Take part in on-call rotations, incident responses, and conduct blameless post-incident evaluations.
• Develop code, manage alerts, refine solutions, and provide support to colleagues.
• Collaborate with stakeholders on requirements analysis and presentations.
• Ensure systems are secure, maintainable, and consistently available.
• Continuously enhance skills through peer reviews and research endeavors.
• Contribute to customer success and align with company objectives.
• Over 5 years of experience in administering Linux systems and related infrastructure in production settings.
• A collaborative mindset as an SRE, familiar with SLIs/SLOs/SLAs, error budgets, blast radius, and conducting blameless postmortems.
• A commitment to automation, minimizing toil, and preventing the recurrence of issues.
• Proven record of creating runbooks that are beneficial for the entire team, not just for individual use.
• Strong foundational knowledge in Kubernetes and its ecosystem.
• Experience with cloud infrastructure, with a preference for AWS; bare-metal knowledge is a plus.
• Proficiency in tool development, particularly with Bash and either Python or Go, or similar languages.
• Familiarity with infrastructure-as-code tools, with a preference for Terraform.
• Experience with CI/CD processes and version control, preferably GitHub.
• Database expertise in one of Postgres, Cassandra, or ClickHouse is preferred.
• Experience managing a production observability stack (metrics, logs, traces), prioritizing meaningful signals over noise.
• Comfortable working with live production infrastructure, possessing strong troubleshooting skills and ownership during incident responses.
• A history of ongoing professional development.
• A self-motivated approach that suits an asynchronous, globally distributed team, with the ability to take on related tasks as needed.
• Flexible working environment – embracing a remote-first culture with coworking options available.
• Generous leave plans – including 4 weeks of paid annual leave, parental leave, birthday leave, and a program for purchasing additional annual leave.
• Health and wellness support – via a wellness allowance and employee wellbeing initiatives.
• Comprehensive learning support – inclusive of a generous study and training allowance alongside 5 days of paid study leave.
• Creative, modern workspaces – designed to inspire when you are not working remotely.
• Motivated, inclusive team – collaborate with industry experts and emerging talent.
• Recognition programs – celebrate achievements with our *Legend* and *Kudos* awards.
Social Discovery Group
Commit
Akamai Technologies
Vetta
Get handpicked remote jobs straight to your inbox weekly.