Senior Reliability Engineer

Posted 5 days ago

This is a fully remote position, open to applicants in Romania.

📋 Description

• Collaborate with application and platform teams to comprehend service architecture, dependencies, critical customer journeys, and reliability risks.

• Establish significant SLIs and SLOs that link technical service behavior with customer and business outcomes.

• Facilitate the adoption of error budgets and align reliability performance with engineering priorities.

• Evaluate services against reliability and production-readiness criteria, including observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery.

• Utilize reliability maturity assessments and scorecards to prioritize impactful enhancements.

• Assist in incident response and analyze complex production issues using logs, metrics, traces, dependencies, and deployment information.

• Review incidents and recurring operational challenges to identify systemic risks and propose long-term improvements.

• Enhance runbooks, alerting, escalation paths, and operational practices.

• Design automation and self-service capabilities to minimize engineering toil.

• Develop tools and automation for reliability workflows, operational readiness, investigation, and remediation.

• Collaborate with Observability Engineering on the necessary telemetry to assess reliability and troubleshoot issues.

• Translate performance tests, capacity assessments, Gamedays, chaos experiments, and failover testing into reliability enhancements.

• Ensure peak-event readiness by assessing service health, SLOs, dependencies, capacity risks, and reliability findings.

• Enhance the reliability of critical customer journeys through monitoring, understanding dependencies, handling failures, and implementing recovery practices.

• Contribute to Reliability Engineering standards, patterns, documentation, and reusable golden paths.

• Integrate reliability capabilities into CI/CD and developer workflows, including SLO-as-Code, telemetry validation, production-readiness checks, and automated operational controls.

• Employ AI-assisted investigation and automation to speed up troubleshooting and lessen manual effort.

• Share expertise and assist engineers in strengthening SRE and production-engineering practices.


⛳️ Requirements

• Solid hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering.

• Strong understanding of core SRE principles, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation.

• Experience in defining or working with SLIs and SLOs for production services.

• Proficient production troubleshooting skills across applications, infrastructure, networks, databases, and service dependencies.

• Good knowledge of distributed systems concepts and reliability patterns such as retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation.

• Familiarity with observability platforms like Datadog or similar, covering logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring.

• Experience in incident response participation, post-incident reviews, and meaningful follow-up actions.

• Practical experience with Kubernetes and cloud infrastructure, preferably AWS.

• Familiarity with infrastructure-as-code tools such as Terraform and modern CI/CD environments.

• Strong automation and software engineering skills, with proficiency in at least one contemporary programming language such as Go, Java, Python, or JavaScript.

• Experience in replacing repetitive operational tasks with automation or self-service.

• Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering.

• Capability to understand application architecture and identify reliability risks across service and infrastructure dependencies.

• Strong analytical and problem-solving abilities.

• Good communication skills and the ability to articulate reliability concepts and recommendations to engineers and engineering leadership.

• Ability to collaborate across multiple teams, balance competing priorities, and drive work to measurable outcomes.

• A mindset focused on automation, continuous improvement, knowledge sharing, and addressing systemic problems.


🏝️ Benefits

• Hybrid & remote working options.

• €1,000 per year for self-development.

• Company share scheme.

• 25 days of annual leave per year.

• 20 days per year to work abroad.

• 5 personal days per year.

• Flexible benefits: travel, sports, hobbies.

• Extended health, dental, and travel insurances.

• Customized well-being programs.

• Career growth sessions.

• Thousands of online courses available through Udemy.

• A variety of engaging office events.

People also viewed

Horizon3.ai12 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH12 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM12 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies12 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)12 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC12 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers