Remotery

Senior Site Reliability Engineer

Posted Jun 26

This is a fully remote position, open to applicants in Spain.

📋 Description

• Oversee the creation of scalable, fault-tolerant, and self-healing systems within a multi-region AWS environment.

• Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to inform architectural choices and error budget strategies.

• Facilitate blameless post-incident analyses to identify systemic root causes and implement long-term preventive solutions.

• Recognize patterns of manual processes and spearhead the development of internal tools and automation to eliminate them permanently.

• Create and uphold automated runbooks and playbooks for routine operational tasks and intricate incident responses.

• Transition from basic monitoring to comprehensive observability, ensuring high cardinality data generates proactive, actionable insights.

• Proactively identify and alleviate operational risks through chaos engineering and architectural assessments.

• Collaborate with software engineers to design systems that prioritize reliability, scalability, and maintainability from the early phases of the SDLC.

• Consistently assess and enhance system performance, capacity, and cost-effectiveness.

• Beyond mere participation, you will enhance the on-call experience to minimize alert fatigue, improve MTTR, and maintain a sustainable rotation health.


⛳️ Requirements

• Bachelor’s degree in Computer Engineering or a related field.

• 5+ years of experience as a Site Reliability Engineer or in a comparable position.

• 3+ years of experience working with AWS services, including a solid understanding of container orchestration.

• 2+ years of hands-on experience with Kubernetes.

• Profound knowledge of observability principles and tools like Prometheus, Datadog, OpenTelemetry, and similar technologies.

• Proven experience in leading incident management and conducting complex postmortem analyses.

• Interest and experience in managing infrastructure as code, particularly using Terraform.

• Familiarity with chaos engineering and other methodologies for testing system resilience.

• Experience with CI/CD tools such as GitHub Actions for automated delivery.

• Proficiency in at least one programming language (e.g., Python, Go, Java) for developing automation and internal tools.

• Knowledge of event-driven architecture (e.g., SNS, SQS).

• Ability to work both independently and collaboratively in a dynamic environment.

• Team-oriented mindset and openness to new concepts.

• Strong communication skills and fluency in English.


🏝️ Benefits

• Remote work opportunities

• Generous paid time off (PTO)

• Wellness allowances

• Learning allowances

• Annual Airalo Away retreat

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers