
Senior Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Romania.
• Collaborate with application and platform teams to comprehend service architecture, dependencies, critical customer journeys, and reliability risks.
• Establish significant SLIs and SLOs that link technical service behavior with customer and business outcomes.
• Facilitate the adoption of error budgets and align reliability performance with engineering priorities.
• Evaluate services against reliability and production-readiness criteria, including observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery.
• Utilize reliability maturity assessments and scorecards to prioritize impactful enhancements.
• Assist in incident response and analyze complex production issues using logs, metrics, traces, dependencies, and deployment information.
• Review incidents and recurring operational challenges to identify systemic risks and propose long-term improvements.
• Enhance runbooks, alerting, escalation paths, and operational practices.
• Design automation and self-service capabilities to minimize engineering toil.
• Develop tools and automation for reliability workflows, operational readiness, investigation, and remediation.
• Collaborate with Observability Engineering on the necessary telemetry to assess reliability and troubleshoot issues.
• Translate performance tests, capacity assessments, Gamedays, chaos experiments, and failover testing into reliability enhancements.
• Ensure peak-event readiness by assessing service health, SLOs, dependencies, capacity risks, and reliability findings.
• Enhance the reliability of critical customer journeys through monitoring, understanding dependencies, handling failures, and implementing recovery practices.
• Contribute to Reliability Engineering standards, patterns, documentation, and reusable golden paths.
• Integrate reliability capabilities into CI/CD and developer workflows, including SLO-as-Code, telemetry validation, production-readiness checks, and automated operational controls.
• Employ AI-assisted investigation and automation to speed up troubleshooting and lessen manual effort.
• Share expertise and assist engineers in strengthening SRE and production-engineering practices.
• Solid hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering.
• Strong understanding of core SRE principles, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation.
• Experience in defining or working with SLIs and SLOs for production services.
• Proficient production troubleshooting skills across applications, infrastructure, networks, databases, and service dependencies.
• Good knowledge of distributed systems concepts and reliability patterns such as retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation.
• Familiarity with observability platforms like Datadog or similar, covering logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring.
• Experience in incident response participation, post-incident reviews, and meaningful follow-up actions.
• Practical experience with Kubernetes and cloud infrastructure, preferably AWS.
• Familiarity with infrastructure-as-code tools such as Terraform and modern CI/CD environments.
• Strong automation and software engineering skills, with proficiency in at least one contemporary programming language such as Go, Java, Python, or JavaScript.
• Experience in replacing repetitive operational tasks with automation or self-service.
• Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering.
• Capability to understand application architecture and identify reliability risks across service and infrastructure dependencies.
• Strong analytical and problem-solving abilities.
• Good communication skills and the ability to articulate reliability concepts and recommendations to engineers and engineering leadership.
• Ability to collaborate across multiple teams, balance competing priorities, and drive work to measurable outcomes.
• A mindset focused on automation, continuous improvement, knowledge sharing, and addressing systemic problems.
• Hybrid & remote working options.
• €1,000 per year for self-development.
• Company share scheme.
• 25 days of annual leave per year.
• 20 days per year to work abroad.
• 5 personal days per year.
• Flexible benefits: travel, sports, hobbies.
• Extended health, dental, and travel insurances.
• Customized well-being programs.
• Career growth sessions.
• Thousands of online courses available through Udemy.
• A variety of engaging office events.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.