
Senior Site Reliability Engineer – AWS, EKS
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in United Kingdom.
• Take charge of the reliability, availability, and operational integrity of production services hosted on AWS and Amazon EKS.
• Manage and troubleshoot Kubernetes clusters in a production setting, encompassing cluster lifecycle, upgrades, node management, networking, scaling, capacity, and workload reliability.
• Engage actively in on-call and pager rotations, accepting responsibility for production incidents.
• Lead or assume a significant technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery, and communication.
• Facilitate technical incident bridges and communicate directly with customers during production escalations as necessary.
• Conduct root cause analysis and lead blameless postmortems, ensuring that incidents lead to tangible engineering enhancements.
• Define, monitor, and enhance SLIs, SLOs, and error budgets for production services.
• Create and maintain actionable alerts, operational runbooks, and automated remediation processes.
• Develop and enhance infrastructure using Terraform/Terragrunt and Infrastructure as Code methodologies.
• Manage GitOps-based delivery environments utilizing tools like Argo CD or FluxCD.
• Optimize Kubernetes scaling and efficiency through technologies such as Karpenter, KEDA, and native Kubernetes autoscaling features.
• Enhance observability using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, and/or ELK.
• Support highly available distributed and event-driven systems, including environments utilizing technologies like Kafka/MSK.
• Design, implement, and validate disaster recovery and business continuity strategies against measurable RTO and RPO goals.
• Identify recurring operational issues and eliminate manual work through automation and engineering solutions.
• Enhance AWS performance, scalability, security, and cost efficiency across production environments.
• Collaborate closely with software, platform, and engineering teams to integrate reliability into systems throughout the development lifecycle.
• Contribute to the ongoing improvement of incident management, operational readiness, and SRE engineering practices.
• Extensive professional experience as a hands-on Site Reliability Engineer, Production Engineer, or senior Platform Engineer with direct production oversight.
• Several years of recent, hands-on experience managing production environments on AWS.
• Strong, demonstrable expertise in operating Amazon EKS in a production context.
• In-depth Kubernetes operational knowledge extending beyond application deployment, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting, and handling production failures.
• Proven experience participating in a production on-call/pager rotation.
• Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, root cause analysis, and post-incident enhancements.
• Practical experience with SLIs, SLOs, error budgets, alerting, and runbooks.
• Strong Infrastructure as Code experience with Terraform and/or Terragrunt.
• Production experience with Kubernetes delivery and GitOps methodologies; experience with Argo CD or FluxCD is highly preferred.
• Solid production observability experience with tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or ELK.
• Experience managing highly available, distributed production systems.
• Strong understanding of AWS networking, IAM, security, availability, and resilience.
• Experience in implementing and testing disaster recovery strategies with measurable RTO/RPO objectives.
• Proven external customer-facing technical experience, including technical discussions, production escalations, architecture/reliability conversations, or incident communication.
• Ability to articulate complex technical issues clearly and make sound decisions during high-pressure production incidents.
• Strong troubleshooting mindset with the capability to work independently during complex production failures.
• Strong professional English communication skills (minimum C1) for regular client interactions.
• A consistent track record demonstrating sustained hands-on production engineering ownership.
• Full-time permanent B2B cooperation.
• Fully remote working environment.
• Senior hands-on engineering position with significant ownership of mission-critical production systems.
• Opportunity to work on complex AWS, Kubernetes, and distributed-system environments at scale.
• Direct influence over reliability engineering, operational practices, and platform enhancements.
• Modern engineering environment with a strong focus on automation, observability, and continuous improvement.
• Collaboration with experienced engineering, platform, and product teams.
• Opportunity to introduce and utilize modern approaches, including AI-assisted engineering and operational automation.
• Long-term opportunities for engineers who wish to remain deeply technical and close to production.
• An inclusive and respectful workplace that values diversity regardless of gender, ethnicity, or background.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.