Senior Site Reliability Engineer

Posted Aug 27

This is a fully remote position, open to applicants in Spain.

📋 Description

• Establish impactful SLIs, SLOs, and reliability benchmarks for the platform.

• Work in tandem with software engineering teams to enhance observability, SLIs, SLOs, and implement reliability best practices.

• Elevate production readiness through effective service ownership, observability, alerting, runbooks, scaling assumptions, rollback strategies, and preparedness for failure modes.

• Enhance the reliability, scalability, and performance of cloud services, Kubernetes, and shared infrastructure.

• Create actionable observability employing metrics, logs, traces, and golden signals utilizing tools like Datadog, Prometheus, and Grafana.

• Enforce operational and security best practices through established guidelines, policies, and automation.

• Minimize alert noise while enhancing the quality of signals.

• Automate repetitive operational tasks using Python or other programming languages.

• Develop self-service Internal Developer Platform features through APIs and Kubernetes operators.

• Increase deployment safety, rollback capabilities, and release observability.

• Boost the reliability of essential stateful systems, including databases, caches, queues, and streaming platforms.

• Engage in on-call duties, troubleshoot issues, coordinate incident responses, and facilitate blameless post-incident reviews.

• Conduct disaster recovery drills and assess cloud/platform usage for cost and resource efficiency improvements.

• Enhance observability, reduce operational burden, improve incident response, establish SRE standards, and empower teams to take ownership of their systems in production.


⛳️ Requirements

• Over 7 years of production experience managing Kubernetes-based platforms and cloud infrastructure.

• Proficient understanding and application of SRE practices, which includes SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.

• Capability to design and enhance observability and alerting for critical systems using metrics, logs, traces, and golden signals.

• Experienced in troubleshooting intricate distributed systems.

• Skillful in writing maintainable software to automate operational tasks and minimize manual intervention.

• Familiar with stateful production systems such as relational databases, caches, queues, or streaming platforms.

• Ability to pragmatically balance reliability, performance, cost, and delivery speed.

• Comfortable working in a transitional environment where SRE practices are being implemented.

• Effective in collaborating with Engineering, Platform, Security, and Product stakeholders.

• Strong communication skills, excellent documentation, and a keen interest in mentoring teams to improve production ownership.

• Demonstrates initiative and accountability, including early identification of risks and driving improvements to completion.

• All applications and CVs must be submitted in English.


🏝️ Benefits

• Remote Flexibility: The option to work from home on days that suit you.

• 25 days of paid vacation per year.

• Intensive working hours in August.

• Comprehensive health, dental, and mental health support via Alan; pre-existing conditions are included.

• €150/month meal allowance provided on your Alan card.

• 50% discount on Ametller Origen prepared meals at the office.

• Flexible remuneration options for additional meal expenses (up to €70/month) and public transport (up to €136/month).

• Provision of a table, ergonomic chair, and monitor for your home office setup.

• Complimentary Spanish language classes.

• Monetary rewards for employee referrals.

• Daily breakfast offered at the office.

• Monthly team-building events.

• A diverse and inclusive international work environment.

People also viewed

Sprezzatura1 day ago

Salesforce Release Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job
Clinician Nexus1 day ago

Manager, DevOps

US flagArizona, +13 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$133.1k – $221.9k/year
ApplyView job
Lenovo1 day ago

CI/CD Engineer

US flagNorth Carolina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$146.2k – $224.1k/year
ApplyView job
MRSOOL | مرسول1 day ago

Site Reliability Engineer II

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CEQUENS1 day ago

DevOps Engineer

EG flagEgypt OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
In All Media1 day ago

DevOps Engineer – Cloud

BR flagBrazil, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers