Senior Site Reliability Engineer – AWS, EKS

Posted Sep 3

This is a fully remote position, open to applicants in Washington.

📋 Description

• Take ownership of the reliability, availability, and operational health of production services hosted on AWS and Amazon EKS.

• Manage and troubleshoot Kubernetes clusters in production, covering lifecycle, upgrades, node management, networking, scaling, capacity, and workload reliability.

• Actively engage in on-call and pager rotations while taking responsibility for production incidents.

• Lead or play a pivotal technical role during P1/P2 and Sev1/Sev2 incidents, focusing on diagnosis, mitigation, recovery, and communication.

• Coordinate technical incident bridges and communicate directly with customers during escalations in production.

• Conduct root cause analysis and facilitate blameless postmortems.

• Define, monitor, and enhance SLIs, SLOs, and error budgets.

• Create and maintain actionable alerts, operational runbooks, and automated remediation processes.

• Develop and enhance infrastructure using Terraform/Terragrunt and Infrastructure as Code methodologies.

• Operate GitOps-based delivery environments utilizing Argo CD or FluxCD.

• Enhance Kubernetes scaling and efficiency through Karpenter, KEDA, and native autoscaling features.

• Develop and improve observability using Prometheus, Grafana, OpenTelemetry, Datadog, and/or ELK.

• Support highly available distributed and event-driven systems, including Kafka/MSK environments.

• Design, implement, and validate disaster recovery and business continuity strategies aligned with RTO and RPO objectives.

• Identify recurring operational challenges and reduce toil through automation and engineering solutions.

• Optimize AWS performance, scalability, security, and cost efficiency.

• Collaborate with software, platform, and engineering teams to enhance reliability throughout the development lifecycle.

• Contribute to the ongoing improvement of incident management, operational readiness, and SRE practices.

• Engage directly with customers during technical discussions and production escalations.


⛳️ Requirements

• Extensive professional experience as a hands-on Site Reliability Engineer, Production Engineer, or senior Platform Engineer with direct ownership of production.

• Several years of recent, hands-on experience managing production environments on AWS.

• Strong, demonstrable experience in operating Amazon EKS in production settings.

• In-depth knowledge of Kubernetes operations, including cluster administration, upgrades, node management, autoscaling, networking, troubleshooting, and scenarios involving production failures.

• Proven experience participating in a production on-call/pager rotation.

• Demonstrated ownership of significant production incidents, encompassing troubleshooting, mitigation, recovery, root cause analysis, and post-incident enhancements.

• Practical experience with SLIs, SLOs, error budgets, alerting, and runbooks.

• Strong expertise in Infrastructure as Code using Terraform and/or Terragrunt.

• Experience with Kubernetes delivery and GitOps practices; familiarity with Argo CD or FluxCD is highly preferred.

• Strong background in production observability with tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or ELK.

• Experience in operating highly available, distributed production systems.

• Comprehensive understanding of AWS networking, IAM, security, availability, and resilience.

• Experience in implementing and testing disaster recovery plans with measurable RTO/RPO objectives.

• Proven experience in external customer-facing technical interactions.

• Ability to articulate complex technical issues clearly and make informed decisions during high-pressure production incidents.

• Strong troubleshooting skills and the capability to work independently during complex production failures.

• Excellent professional English communication skills (minimum C1) for regular client interactions.

• A solid track record showcasing sustained hands-on production engineering ownership.


🏝️ Benefits

• Full-time permanent B2B cooperation.

• Fully remote working environment.

• Senior hands-on engineering role with significant ownership of business-critical production systems.

• Opportunity to engage with complex AWS, Kubernetes, and distributed-system environments at scale.

• Direct impact on reliability engineering, operational practices, and platform enhancements.

• Modern engineering environment emphasizing automation, observability, and continuous improvement.

• Collaboration with experienced engineering, platform, and product teams.

• Opportunity to implement and utilize modern approaches, including AI-assisted engineering and operational automation.

• Long-term opportunity for engineers aiming to remain deeply technical and close to production.

• An inclusive and respectful workplace, regardless of gender, ethnicity, or background.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF14 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam18 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers