Senior Site Reliability Engineer, AWS Multi-region

Posted 17 hours ago

This is a fully remote position, open to applicants in Argentina, +2 more countries.

📋 Description

• Facilitate resilience across multiple regions and oversee both automated and manual failover of essential services utilized throughout the organization.

• Take ownership of around 10–15 services and lead pilot initiatives on 4–5 applications to assess the effectiveness of the multi-region strategy (pilot-light vs active-active).

• Design foundational services with a multi-region priority and eliminate dependencies between regions.

• Implement data replication across regions with minimal latency.

• Set up seamless traffic failover mechanisms.

• Develop and maintain comprehensive runbooks and operational playbooks.

• Automate the deployment processes for multi-region environments.

• Conduct monitoring and perform cost evaluations.

• Deliver a thoroughly validated multi-region architecture.

• Create automation templates for consistent deployment practices.

• Successfully execute at least one failover test.

• Design, implement, and manage highly available cloud infrastructures.

• Transition production workloads from single-region setups to multi-region configurations.

• Enhance disaster recovery features.

• Build resilient infrastructure using automation and Infrastructure as Code principles.

• Collaborate effectively with engineering teams and business partners.


⛳️ Requirements

• Demonstrated experience in designing and implementing AWS migrations from single-region to multi-region, including strategies for regional failover and traffic management (Route 53, AWS Global Accelerator).

• Extensive knowledge in Infrastructure as Code (Terraform, CloudFormation, or CDK).

• Proficiency in containerization and orchestration technologies (EKS, ECS, Kubernetes).

• Experience with CI/CD processes for multi-region deployments.

• Practical experience with cross-region data replication techniques.

• Familiarity with observability practices (metrics, logs, tracing, alerting).

• Strong troubleshooting skills in large-scale production environments.

• Capacity for technical ownership.

• Ability to collaborate effectively with engineering teams and business stakeholders.


🏝️ Benefits

• Paid time off or vacation days.

• Paid holidays, contingent on the client and country.

• Opportunities for professional development and learning.

• Continuous support from our People Experience team.

• Initiatives for recognition and internal engagement.

• Flexibility for remote work.

• Opportunities to collaborate with teams globally.

• Access to tools and resources that enhance your work.

• Potential for career advancement based on performance, business requirements, and available positions.

People also viewed

Karat4 hours ago

Senior DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM17 hours ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM17 hours ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Charger Logistics Inc.1 day ago

Site Reliability Engineer

IN flagIndia OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dev.Pro1 day ago

Senior DevOps Engineer

BR flagBrazil, +2 more countriesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
EVT1 day ago

Senior DevSecOps Engineer

BR flagBrazil OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers