
Site Reliability Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Costa Rica.
• Design and implement automation solutions to minimize manual and repetitive operational tasks in productive environments.
• Implement and evolve observability practices using metrics, logs, and traces.
• Define and maintain SLIs and SLOs in collaboration with development teams and business stakeholders.
• Participate in incident response and on-call schemes, performing diagnostics, containment, and service recovery.
• Lead and document post-incident analyses and Root Cause Analysis (RCA), identifying areas for improvement and preventive actions.
• Promote the "you build it, you run it" model.
• Design and maintain change control mechanisms and deployments integrated into CI/CD pipelines.
• Analyze capacity, performance, and behavior under load of services, proposing architectural or configuration enhancements.
• Evaluate and strengthen the resilience, availability, and recovery from failures of platforms and services.
• Document standards, procedures, runbooks, and best practices for Site Reliability Engineering.
• Work closely with development and platform teams.
• Professional experience in Site Reliability Engineering, DevOps, Cloud, Infrastructure, Systems Engineering, or Operations.
• Hands-on experience working with critical services and productive environments.
• Strong knowledge of Linux and troubleshooting infrastructure and applications.
• Experience with cloud computing platforms.
• Experience in monitoring and observability using metrics, logs, traces, and alerting.
• Experience in handling production incidents, troubleshooting, and Root Cause Analysis.
• Understanding of SLI, SLO, and SLA concepts and their application in reliability management.
• Experience with automation and scripting.
• Knowledge of CI/CD and modern deployment processes.
• Experience with Infrastructure as Code, ideally Terraform.
• Knowledge of high availability, resilience, disaster recovery, and scalability principles.
• Ability to collaborate effectively with development, infrastructure, and operations teams.
• Preferred experience with AWS and cloud-native architectures.
• Preferred experience with Terraform and infrastructure automation.
• Preferred experience with Docker and Kubernetes.
• Preferred knowledge of observability tools such as Prometheus, Grafana, Datadog, Splunk, CloudWatch, or equivalents.
• Preferred experience working with High Availability and Disaster Recovery strategies.
• Preferred experience in capacity, performance, and scalability analysis.
• Preferred experience participating in on-call schemes and managing critical incidents.
• Desired certifications related to AWS, Cloud Operations, or Site Reliability Engineering.
• Private health insurance.
• Solidarity association.
• Vacation and holidays in accordance with Costa Rican legislation.
• Opportunities for learning and professional growth.
• Participation in high-impact technological projects.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.