Senior Site Reliability Engineer

Posted Sep 2

This is a fully remote position, open to applicants in Argentina.

📋 Description

• Define, implement, and enhance SLIs, SLOs, and error budgets for essential applications and platform services.

• Collaborate with application teams to boost reliability, fault tolerance, scalability, and operational readiness.

• Identify and resolve recurring reliability challenges through root cause analysis, automation, and architectural enhancements.

• Design robust systems to withstand Azure region, zone, network, dependency, and deployment failures.

• Engage in production readiness assessments for services, releases, and infrastructure modifications.

• Establish observability across applications, infrastructure, networks, and cloud services.

• Implement monitoring for latency, traffic, errors, and saturation levels.

• Create dashboards, alerts, logs, traces, and metrics utilizing observability and APM platforms.

• Develop service health dashboards for engineering, operations, and leadership teams.

• Analyze performance metrics, bottlenecks, saturation trends, and capacity risks.

• Enhance backup, disaster recovery, failover, and business continuity practices.

• Apply resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling.

• Support and enhance production workloads operating on Microsoft Azure.

• Collaborate on secure and scalable Azure architecture.

• Enforce Azure operational standards regarding tagging, monitoring, backup, recovery, identity, security, and cost awareness.

• Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention strategies.

• Refine incident response processes and runbooks in coordination with Incident Management, NOC, Help Desk, and application teams.

• Support cloud security and compliance controls related to identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.

• Collaborate with software engineering, product, security, and operations leadership.

• Mentor the SRE team and act as a technical authority for observability and reliability across mission-critical healthcare workloads.


⛳️ Requirements

• Over 7 years of experience in site reliability, including hands-on ownership of mission-critical services, derived from a mix of relevant work experience, training, military service, or education.

• Extensive Microsoft Azure expertise, including Azure Monitor, Application Insights, AKS, and cloud-native operations across hybrid infrastructures.

• Demonstrated experience in designing and implementing an SLO program with SLIs, SLOs, and error budget policies in production environments.

• Practical experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.

• Hands-on experience configuring and implementing Datadog for monitoring, observability, and alerting purposes.

• Identity and credential verification, live or video interviews, and screening for fraud or misrepresentation.

• Preferred: experience in establishing or leading an SRE practice across multiple engineering teams.

• Preferred: strong incident command experience and conducting blameless postmortems.

• Preferred: background in healthcare IT and familiarity with HIPAA, HITRUST, or similar compliance frameworks.

• Preferred: multi-cloud reliability experience, particularly with AWS alongside Azure.

• Preferred: experience in chaos engineering and resiliency testing.

• Preferred: expertise in Infrastructure as Code using Terraform, Bicep, or Ansible.

• Preferred: proficiency in scripting and programming languages such as Python, PowerShell, or Go.

• Preferred: experience with security controls, vulnerability management, compliance audits, and cloud governance.

• Preferred: possession of recognized industry certifications such as Azure Solutions Architect, Google SRE certificate, or CKA.


🏝️ Benefits

• Inclusive benefits program focused on employees and their families.

• Customized programs addressing the unique needs of employees.

• Significant opportunities for career growth, leadership, and development.

• An inclusive workplace culture that fosters innovation.

• Resources for candidates and support from recruiters.

People also viewed

Probis18 hours ago

Senior DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Entarian20 hours ago

DevSecOps Engineer – Mid

US flagCalifornia, +2 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Velera21 hours ago

Manager, Technology Delivery – ADO/DevOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$117k – $152.1k/year
ApplyView job
RR Donnelley22 hours ago

Senior Linux, Infrastructure Automation Engineer

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$91.8k – $146.8k/year
ApplyView job
NVIDIA23 hours ago

Engineering Manager, Reliability Engineering – EDA Infrastructure

US flagCalifornia, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$224k – $431.3k/year
ApplyView job
Vital Tech Solutions1 day ago

Senior Site Reliability Engineer, SRE

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers