Remotery

Senior Site Reliability Engineer – Remote within EMEA

Posted Aug 12

This is a fully remote position, open to applicants in Poland, +4 more states.

📋 Description

• Design and oversee a highly available, secure, and scalable infrastructure across various AWS accounts and environments utilizing infrastructure as code.

• Develop and manage Amazon EKS clusters, focusing on networking policies, persistent storage, and scaling strategies.

• Take ownership of cloud networking, Kubernetes, databases, and messaging systems.

• Lead extensive automation initiatives and establish standards for Terraform/Terragrunt and GitOps (ArgoCD).

• Promote automation adoption to minimize manual operational tasks and ensure consistent, repeatable environments.

• Address intricate, cross-service infrastructure challenges.

• Enhance reliability and observability through tools like Grafana, Prometheus, Loki, Tempo, and Mimir.

• Engage in on-call rotations and spearhead production incident responses.

• Compose runbooks, Architectural Decision Records (ADRs), and postmortems.

• Direct team security initiatives related to IAM, encryption, and secure logging practices.

• Examine infrastructure spending and drive cost optimization through rightsizing, autoscaling, and FinOps methodologies.

• Mentor mid-level and junior SREs, offering feedback and aiding in onboarding processes.

• Collaborate with product engineering squads and represent Platform Infrastructure in cross-team projects.

• Suggest designs, evaluate engineering efforts, and facilitate build-vs-buy and cost/reliability assessments.


⛳️ Requirements

• Strong background in Linux systems administration and proficiency in scripting with Python.

• Extensive knowledge of AWS services like EKS, IAM, VPC networking, RDS, S3, SQS, and the Well-Architected Framework.

• Experience in multi-account AWS environments is highly advantageous.

• Hands-on experience with production-scale Kubernetes/EKS operations and troubleshooting.

• Development of Helm charts, CNI networking, Cilium, pod networking/IPAM, and container security expertise.

• Proficient in Terraform, with a preference for Terragrunt experience.

• Familiarity with GitOps practices using ArgoCD.

• Production-scale experience with database systems such as Postgres, MySQL, and/or MongoDB, including Aurora.

• Experience with CI/CD tools like Jenkins and/or GitHub Actions.

• Understanding of blue-green and canary deployment methodologies.

• Experience maintaining the Grafana observability stack, comprising Grafana, Prometheus, Loki, Tempo, and Mimir.

• Practical experience in incident management, including on-call rotations, structured incident response, runbooks, and postmortems.

• Familiarity with 12-Factor App principles and cost optimization/FinOps strategies.

• A methodical, data-driven approach to troubleshooting.

• Excellent written communication skills.

• Capability to drive initiatives with unclear ownership and take responsibility for results.

• Proven track record of mentoring less experienced engineers and providing direct, constructive feedback.

• Several years of hands-on experience in production infrastructure/SRE at a senior individual contributor level.

• Nice to have: Experience with GCP, Kafka/AWS MSK, regulated or compliance-sensitive environments, and platform/DevEx roadmap experience.


🏝️ Benefits

• A global, inclusive team environment.

• Flexibility to work remotely or from a modern and welcoming office in Riga.

• Flexible working hours (start your day no later than 11 AM).

• Private health insurance coverage.

• Two additional paid days off to prioritize mental or physical well-being.

• One extra paid day off to celebrate your Birthday or any other personal celebration of your choice.

• Opportunities for internal and external learning.

• Access to mentorship, internal meetups, and hackathons, both on-site and online.

• Complimentary healthy lunch for those working from the Rīga office.

• Ability to design and order your own merchandise via our platforms with employee discounts.

• Team-building events and social gatherings.

People also viewed

CWILL16 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3717 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT18 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group18 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo18 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch18 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers