
Staff Software Engineer – Databases SRE
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in United Kingdom.
• Assist Grafana Cloud's most valuable customers by ensuring the reliability of their databases.
• Collaborate closely with product engineering teams.
• Take ownership of production reliability for environments with high service level agreements (SLAs).
• Design and implement automated processes to enhance reliability practices.
• Ensure that customers achieve their service level objective (SLO) targets.
• Lead incident response efforts and conduct reviews.
• Contribute to the creation of design documentation and participate in code reviews.
• Develop automation solutions to minimize repetitive tasks.
• Enhance the quality of alerts and reduce the number of escalations.
• Over 8 years of engineering experience, with at least 4 years in Site Reliability Engineering (SRE), Cloud Reliability Engineering (CRE), or production engineering.
• Proficient in Kubernetes, particularly within AWS, GCP, or Azure environments.
• Familiarity with infrastructure-as-code tools such as Helm, Terraform, Jsonnet, etc.
• Proven experience in technical leadership roles.
• Experience in operating multi-tenant systems within production settings.
• Strong background in designing and implementing service level objectives (SLOs).
• Proficiency in one or more programming languages (e.g., Go, Python, Java).
• Knowledge of Linux internals.
• Exceptional problem-solving abilities.
• Experience in incident response and conducting post-incident reviews.
• Ability to analyze performance, scalability, and potential failure modes.
• Capable of working autonomously within an engineering team.
• Equity.
• Bonus (if applicable).
• 30 days of annual leave.
• Grafana Shutdown Days.
• In-person onboarding.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.