Remotery

Site Reliability Engineer – Azure, Observability, Scripting

Posted Aug 5

This is a fully remote position, open to applicants in Colombia, +4 more states.

📋 Description

• Ensure the dependability, scalability, and performance of cloud-native platforms operating on Microsoft Azure and Kubernetes.

• Enhance system availability, monitoring, incident response, and infrastructure automation within production settings.

• Design and implement observability solutions, which include monitoring, metrics, exporters, alerting rules, and dashboards.

• Define and establish service level indicators, objectives, and error budgets.

• Create alerting and incident response processes, author runbooks, and support on-call practices.

• Design and test backup, restore, and disaster recovery solutions in line with RPO and RTO targets.

• Troubleshoot Kubernetes workloads, manage resources, and carry out cluster upgrades.

• Read and modify infrastructure as code and Azure DevOps pipelines.

• Promote operational excellence through automation and scripting.


⛳️ Requirements

• A minimum of 6 years of relevant experience.

• At least 3 years of experience managing Kubernetes in a production environment.

• Experience in site reliability engineering or production operations for Kubernetes workloads at scale.

• Proficiency with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy.

• Familiarity with Prometheus and Grafana, covering metrics, exporters, recording and alerting rules, and dashboard design.

• Expertise in defining and implementing service level indicators, objectives, and error budgets.

• Experience in designing alerting and incident response, which includes runbook authoring and on-call practices.

• Background in designing and testing backup, restore, and disaster recovery solutions, including validation against RPO and RTO targets.

• Kubernetes operations experience, encompassing workload troubleshooting, resource management, and cluster upgrades.

• Capability to read and modify infrastructure as code using Terraform or Bicep and Azure DevOps pipelines.

• Scripting proficiency in Python, PowerShell, or Bash.

• Professional proficiency in English.


🏝️ Benefits

• Remote work opportunities.

• 13 floating holidays.

• 15 vacation days per year upon completion.

• Positive working environment.

• Equal opportunity employment without discrimination based on protected characteristics.

People also viewed

CWILL11 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3712 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT13 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group13 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo14 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch14 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers