
Senior Site Reliability Engineer
Posted Jun 26

Posted Jun 26
This is a fully remote position, open to applicants in Spain.
• Oversee the creation of scalable, fault-tolerant, and self-healing systems within a multi-region AWS environment.
• Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to inform architectural choices and error budget strategies.
• Facilitate blameless post-incident analyses to identify systemic root causes and implement long-term preventive solutions.
• Recognize patterns of manual processes and spearhead the development of internal tools and automation to eliminate them permanently.
• Create and uphold automated runbooks and playbooks for routine operational tasks and intricate incident responses.
• Transition from basic monitoring to comprehensive observability, ensuring high cardinality data generates proactive, actionable insights.
• Proactively identify and alleviate operational risks through chaos engineering and architectural assessments.
• Collaborate with software engineers to design systems that prioritize reliability, scalability, and maintainability from the early phases of the SDLC.
• Consistently assess and enhance system performance, capacity, and cost-effectiveness.
• Beyond mere participation, you will enhance the on-call experience to minimize alert fatigue, improve MTTR, and maintain a sustainable rotation health.
• Bachelor’s degree in Computer Engineering or a related field.
• 5+ years of experience as a Site Reliability Engineer or in a comparable position.
• 3+ years of experience working with AWS services, including a solid understanding of container orchestration.
• 2+ years of hands-on experience with Kubernetes.
• Profound knowledge of observability principles and tools like Prometheus, Datadog, OpenTelemetry, and similar technologies.
• Proven experience in leading incident management and conducting complex postmortem analyses.
• Interest and experience in managing infrastructure as code, particularly using Terraform.
• Familiarity with chaos engineering and other methodologies for testing system resilience.
• Experience with CI/CD tools such as GitHub Actions for automated delivery.
• Proficiency in at least one programming language (e.g., Python, Go, Java) for developing automation and internal tools.
• Knowledge of event-driven architecture (e.g., SNS, SQS).
• Ability to work both independently and collaboratively in a dynamic environment.
• Team-oriented mindset and openness to new concepts.
• Strong communication skills and fluency in English.
• Remote work opportunities
• Generous paid time off (PTO)
• Wellness allowances
• Learning allowances
• Annual Airalo Away retreat
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.