
Site Reliability Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Argentina.
• Design, implement, and sustain observability solutions across Azure.
• Establish and standardize SLIs/SLOs/SLAs to evaluate service health.
• Create dashboards and automated alerts to detect service degradations.
• Develop and maintain on-call runbooks and playbooks.
• Facilitate post-incident “blameless” retrospectives.
• Create automation for self-healing systems and monitoring remediation.
• Contribute to disaster recovery and business continuity strategies.
• Collaborate with teams in security, networking, and system engineering.
• Guide engineering staff in observability tools and incident management.
• 3-5 years of experience in Site Reliability Engineering.
• Proficiency in Azure (VMs, AKS, Application Insights, Azure Monitor).
• Practical experience with enterprise observability tools (Datadog, Dynatrace, Grafana).
• Strong understanding of metrics, logs, traces, and monitoring of distributed systems.
• Expertise in Terraform, Bicep, Ansible, or comparable tools.
• Familiarity with AKS logging in containerized environments.
• Experience with AI-driven monitoring or predictive alerting technologies.
• Scripting capabilities in Python or PowerShell.
• Experience with resiliency testing tools like Gremlin or Chaos Mesh.
• Health insurance.
• Flexible working hours.
• Opportunities for professional development.
• Paid time off.
• Options for remote work.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.