
Senior Site Reliability Engineer
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in Argentina.
• Define, implement, and enhance SLIs, SLOs, and error budgets for essential applications and platform services.
• Collaborate with application teams to boost reliability, fault tolerance, scalability, and operational readiness.
• Identify and resolve recurring reliability challenges through root cause analysis, automation, and architectural enhancements.
• Design robust systems to withstand Azure region, zone, network, dependency, and deployment failures.
• Engage in production readiness assessments for services, releases, and infrastructure modifications.
• Establish observability across applications, infrastructure, networks, and cloud services.
• Implement monitoring for latency, traffic, errors, and saturation levels.
• Create dashboards, alerts, logs, traces, and metrics utilizing observability and APM platforms.
• Develop service health dashboards for engineering, operations, and leadership teams.
• Analyze performance metrics, bottlenecks, saturation trends, and capacity risks.
• Enhance backup, disaster recovery, failover, and business continuity practices.
• Apply resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling.
• Support and enhance production workloads operating on Microsoft Azure.
• Collaborate on secure and scalable Azure architecture.
• Enforce Azure operational standards regarding tagging, monitoring, backup, recovery, identity, security, and cost awareness.
• Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention strategies.
• Refine incident response processes and runbooks in coordination with Incident Management, NOC, Help Desk, and application teams.
• Support cloud security and compliance controls related to identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.
• Collaborate with software engineering, product, security, and operations leadership.
• Mentor the SRE team and act as a technical authority for observability and reliability across mission-critical healthcare workloads.
• Over 7 years of experience in site reliability, including hands-on ownership of mission-critical services, derived from a mix of relevant work experience, training, military service, or education.
• Extensive Microsoft Azure expertise, including Azure Monitor, Application Insights, AKS, and cloud-native operations across hybrid infrastructures.
• Demonstrated experience in designing and implementing an SLO program with SLIs, SLOs, and error budget policies in production environments.
• Practical experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.
• Hands-on experience configuring and implementing Datadog for monitoring, observability, and alerting purposes.
• Identity and credential verification, live or video interviews, and screening for fraud or misrepresentation.
• Preferred: experience in establishing or leading an SRE practice across multiple engineering teams.
• Preferred: strong incident command experience and conducting blameless postmortems.
• Preferred: background in healthcare IT and familiarity with HIPAA, HITRUST, or similar compliance frameworks.
• Preferred: multi-cloud reliability experience, particularly with AWS alongside Azure.
• Preferred: experience in chaos engineering and resiliency testing.
• Preferred: expertise in Infrastructure as Code using Terraform, Bicep, or Ansible.
• Preferred: proficiency in scripting and programming languages such as Python, PowerShell, or Go.
• Preferred: experience with security controls, vulnerability management, compliance audits, and cloud governance.
• Preferred: possession of recognized industry certifications such as Azure Solutions Architect, Google SRE certificate, or CKA.
• Inclusive benefits program focused on employees and their families.
• Customized programs addressing the unique needs of employees.
• Significant opportunities for career growth, leadership, and development.
• An inclusive workplace culture that fosters innovation.
• Resources for candidates and support from recruiters.
Probis
Entarian
Velera
RR Donnelley
Get handpicked remote jobs straight to your inbox weekly.