
Senior Site Reliability Engineer
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Spain.
• Establish impactful SLIs, SLOs, and reliability benchmarks for the platform.
• Work in tandem with software engineering teams to enhance observability, SLIs, SLOs, and implement reliability best practices.
• Elevate production readiness through effective service ownership, observability, alerting, runbooks, scaling assumptions, rollback strategies, and preparedness for failure modes.
• Enhance the reliability, scalability, and performance of cloud services, Kubernetes, and shared infrastructure.
• Create actionable observability employing metrics, logs, traces, and golden signals utilizing tools like Datadog, Prometheus, and Grafana.
• Enforce operational and security best practices through established guidelines, policies, and automation.
• Minimize alert noise while enhancing the quality of signals.
• Automate repetitive operational tasks using Python or other programming languages.
• Develop self-service Internal Developer Platform features through APIs and Kubernetes operators.
• Increase deployment safety, rollback capabilities, and release observability.
• Boost the reliability of essential stateful systems, including databases, caches, queues, and streaming platforms.
• Engage in on-call duties, troubleshoot issues, coordinate incident responses, and facilitate blameless post-incident reviews.
• Conduct disaster recovery drills and assess cloud/platform usage for cost and resource efficiency improvements.
• Enhance observability, reduce operational burden, improve incident response, establish SRE standards, and empower teams to take ownership of their systems in production.
• Over 7 years of production experience managing Kubernetes-based platforms and cloud infrastructure.
• Proficient understanding and application of SRE practices, which includes SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
• Capability to design and enhance observability and alerting for critical systems using metrics, logs, traces, and golden signals.
• Experienced in troubleshooting intricate distributed systems.
• Skillful in writing maintainable software to automate operational tasks and minimize manual intervention.
• Familiar with stateful production systems such as relational databases, caches, queues, or streaming platforms.
• Ability to pragmatically balance reliability, performance, cost, and delivery speed.
• Comfortable working in a transitional environment where SRE practices are being implemented.
• Effective in collaborating with Engineering, Platform, Security, and Product stakeholders.
• Strong communication skills, excellent documentation, and a keen interest in mentoring teams to improve production ownership.
• Demonstrates initiative and accountability, including early identification of risks and driving improvements to completion.
• All applications and CVs must be submitted in English.
• Remote Flexibility: The option to work from home on days that suit you.
• 25 days of paid vacation per year.
• Intensive working hours in August.
• Comprehensive health, dental, and mental health support via Alan; pre-existing conditions are included.
• €150/month meal allowance provided on your Alan card.
• 50% discount on Ametller Origen prepared meals at the office.
• Flexible remuneration options for additional meal expenses (up to €70/month) and public transport (up to €136/month).
• Provision of a table, ergonomic chair, and monitor for your home office setup.
• Complimentary Spanish language classes.
• Monetary rewards for employee referrals.
• Daily breakfast offered at the office.
• Monthly team-building events.
• A diverse and inclusive international work environment.
Sprezzatura
Clinician Nexus
Lenovo
MRSOOL | مرسول
Get handpicked remote jobs straight to your inbox weekly.