
Site Reliability Engineer – Azure, Observability, Scripting
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Colombia, +4 more states.
• Ensure the dependability, scalability, and performance of cloud-native platforms operating on Microsoft Azure and Kubernetes.
• Enhance system availability, monitoring, incident response, and infrastructure automation within production settings.
• Design and implement observability solutions, which include monitoring, metrics, exporters, alerting rules, and dashboards.
• Define and establish service level indicators, objectives, and error budgets.
• Create alerting and incident response processes, author runbooks, and support on-call practices.
• Design and test backup, restore, and disaster recovery solutions in line with RPO and RTO targets.
• Troubleshoot Kubernetes workloads, manage resources, and carry out cluster upgrades.
• Read and modify infrastructure as code and Azure DevOps pipelines.
• Promote operational excellence through automation and scripting.
• A minimum of 6 years of relevant experience.
• At least 3 years of experience managing Kubernetes in a production environment.
• Experience in site reliability engineering or production operations for Kubernetes workloads at scale.
• Proficiency with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy.
• Familiarity with Prometheus and Grafana, covering metrics, exporters, recording and alerting rules, and dashboard design.
• Expertise in defining and implementing service level indicators, objectives, and error budgets.
• Experience in designing alerting and incident response, which includes runbook authoring and on-call practices.
• Background in designing and testing backup, restore, and disaster recovery solutions, including validation against RPO and RTO targets.
• Kubernetes operations experience, encompassing workload troubleshooting, resource management, and cluster upgrades.
• Capability to read and modify infrastructure as code using Terraform or Bicep and Azure DevOps pipelines.
• Scripting proficiency in Python, PowerShell, or Bash.
• Professional proficiency in English.
• Remote work opportunities.
• 13 floating holidays.
• 15 vacation days per year upon completion.
• Positive working environment.
• Equal opportunity employment without discrimination based on protected characteristics.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.