Remotery

SRE, Site Reliability Engineering

Posted Jul 18

This is a fully remote position, open to applicants in United States.

📋 Description

• Tailor observability requirements to fit each technical solution, ensuring comprehensive coverage, visibility, and operational effectiveness.

• Set up and manage dashboards, metrics, alerts, and essential business controls.

• Assess solution robustness through chaos testing and scalability evaluations under load conditions.

• Apply resilient design patterns, including circuit breakers, fallbacks, and retries within distributed architectures.

• Detect and automate manual processes utilizing infrastructure-as-code tools to minimize MTTR.

• Spearhead the establishment of self-remediation workflows while fostering a culture of continuous improvement in operations.

• Work collaboratively with development and architecture teams to guarantee technical excellence across important user journeys.


⛳️ Requirements

• At least 3 years of experience in leading technology resilience and observability within complex environments.

• Demonstrated expertise in automating operational tasks and managing incidents following SRE/DevOps practices.

• Proficient in observability tools: Dynatrace (primary hands-on), Grafana, Prometheus, OpenTelemetry, and the ELK Stack.

• Skilled in automation and IaC tools: Ansible, Terraform, Terragrunt, and Monaco (Monitoring as Code).

• Experienced in containerization technologies: Kubernetes (AKS, EKS), OpenShift (advanced level), and Docker.

• Knowledgeable in programming languages: Python (advanced), Bash, YAML, and PowerShell.

• Familiar with Cloud & Infrastructure platforms: Azure, AWS, or GCP (Networking, Security, and Compute).

• Expertise in reliability management, including the establishment of SLIs, SLOs, SLAs, and Error Budget management.

• Proficient in CI/CD practices: Git, Jenkins, Azure DevOps, and GitHub Actions.

• Experienced in resilience engineering techniques: Chaos Engineering, circuit breaker patterns, and Canary/Blue-Green deployments.


🏝️ Benefits

• Technical and personal challenges that will facilitate your continuous growth.

• A supportive team dedicated to your physical and mental wellness.

• A vibrant, collaborative culture of continuous improvement with ample learning opportunities and a community ready to assist you.

• KaizenHub, a program aimed at enhancing your skills, providing feedback, mentoring, and coaching through Sofka U.

• Initiatives like Happy Kaizen and WeSofka that promote your physical and emotional wellbeing.

People also viewed

The Codest3 days ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUM3 days ago

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
Sólides3 days ago

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Resilinc3 days ago

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group3 days ago

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLER3 days ago

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers