Senior Site Reliability Engineer, AWS, EKS

Posted Sep 3

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Take ownership of the reliability, availability, and operational health of production services hosted on AWS and Amazon EKS.

• Manage and resolve issues within Kubernetes clusters in production, covering lifecycle management, upgrades, nodes, networking, scaling, as well as capacity and workload reliability.

• Actively engage in on-call and pager rotations, taking responsibility for production incidents.

• Lead or play a crucial technical role during P1/P2 and Sev1/Sev2 incidents, which includes diagnosis, mitigation, recovery, and effective communication.

• Coordinate technical incident bridges and communicate directly with customers during production escalations.

• Conduct root cause analysis and facilitate blameless postmortems.

• Establish, monitor, and enhance SLIs, SLOs, and error budgets.

• Create and maintain actionable alerts, operational runbooks, and automated remediation processes.

• Build and enhance infrastructure utilizing Terraform/Terragrunt and Infrastructure as Code methodologies.

• Operate GitOps-based delivery environments through Argo CD or FluxCD.

• Enhance Kubernetes scaling and efficiency using Karpenter, KEDA, and native Kubernetes autoscaling features.

• Develop and improve observability frameworks using Prometheus, Grafana, OpenTelemetry, Datadog, and/or ELK.

• Support highly available distributed and event-driven systems, including Kafka/MSK environments.

• Design, implement, and validate disaster recovery and business continuity strategies against RTO and RPO objectives.

• Remove recurring operational issues and toil through automation and engineering solutions.

• Optimize AWS performance, scalability, security, and cost-effectiveness.

• Collaborate with software, platform, and engineering teams to ensure reliability throughout the development lifecycle.

• Contribute to incident management, operational readiness, and SRE engineering practices.


⛳️ Requirements

• Significant hands-on experience as a Site Reliability Engineer, Production Engineer, or senior Platform Engineer with direct production responsibilities.

• Several years of recent, hands-on experience managing production environments on AWS.

• Strong, demonstrable expertise in operating Amazon EKS in a production setting.

• In-depth knowledge of Kubernetes operations, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting, and addressing production failure scenarios.

• Proven experience in participating in a production on-call/pager rotation.

• Demonstrated ownership of significant production incidents, including troubleshooting, mitigation, recovery, root cause analysis, and post-incident enhancements.

• Practical knowledge of SLIs, SLOs, error budgets, alerting, and runbooks.

• Strong experience with Infrastructure as Code using Terraform and/or Terragrunt.

• Production experience with Kubernetes delivery and GitOps practices; proficiency in Argo CD or FluxCD is strongly preferred.

• Robust experience in production observability with Prometheus, Grafana, OpenTelemetry, Datadog, or ELK.

• Experience managing highly available, distributed production systems.

• Strong understanding of AWS networking, IAM, security, availability, and resilience.

• Experience in implementing and testing disaster recovery strategies with measurable RTO/RPO objectives.

• Proven technical experience engaging with external customers.

• Ability to clearly articulate complex technical challenges and make sound decisions during high-pressure production incidents.

• Strong troubleshooting mindset and capability to work independently during intricate production failures.

• Excellent professional English communication skills (minimum C1) for regular client interactions.

• A solid track record demonstrating consistent hands-on production engineering ownership.


🏝️ Benefits

• Full-time permanent B2B collaboration.

• Completely remote working environment.

• Senior hands-on engineering role with significant ownership of critical production systems.

• Opportunity to work with complex AWS, Kubernetes, and distributed system environments at scale.

• Direct influence over reliability engineering, operational practices, and platform enhancements.

• Modern engineering environment with a strong focus on automation, observability, and continuous improvement.

• Collaborate with seasoned engineering, platform, and product teams.

• Opportunity to introduce and utilize modern methodologies, including AI-assisted engineering and operational automation.

• Long-term opportunities for engineers who wish to remain deeply technical and close to production.

• An inclusive and respectful workplace that values diversity in gender, ethnicity, and background.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers