Remotery

Director of SRE

Posted Aug 4

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of and implement the SRE strategy along with a multi-quarter roadmap focusing on reliability, observability, incident management, QA maturity, and release engineering.

• Define, assess, and enhance SLAs, SLOs, error budgets, uptime, performance, and operational health metrics.

• Direct efforts towards production reliability, which includes monitoring, alerting, on-call operations, incident response, root cause analysis, and reducing MTTR.

• Establish standards for release readiness, safety controls for deployment, and quality gates.

• Oversee external SRE vendors and partners, managing service delivery, SLA governance, escalations, performance evaluations, and compliance expectations.

• Lead the QA engineering strategy with a focus on automation, prevention of regression, test coverage, and minimizing defects in production.

• Collaborate with Security and Engineering leaders on cloud infrastructure, CI/CD pipelines, operational tooling, HIPAA, SOC2, and internal security protocols.

• Supervise Azure AKS, Kubernetes, GitOps workflows, CI/CD pipelines, GitHub Actions, secrets management, access controls, and audit readiness.

• Promote the maturity of observability through tools such as Grafana, Prometheus, logging platforms, tracing tools, and automated alerting frameworks.

• Work alongside Product, Platform, and Engineering teams to integrate reliability and quality practices throughout the software development lifecycle.

• Build, mentor, and expand the SRE and QA teams.

• Propel AI-enabled automation and intelligent tools to reduce manual toil and enhance operational excellence.


⛳️ Requirements

• A minimum of 12 years of experience in SRE, infrastructure, or platform engineering.

• At least 5 years of experience in engineering leadership roles.

• Demonstrated ownership of site reliability for complex, multi-tenant SaaS platforms with demanding availability standards.

• Proven experience in defining SLA and SLO frameworks, error budgets, and incident management processes at scale.

• Experience managing infrastructure or SRE service vendors, including SLA governance and performance management.

• Background in leading QA or quality engineering functions, test automation maturity, and ownership of release gates.

• Strong hands-on experience with Microsoft Azure, ideally including AKS, networking, storage, IAM, and security services.

• In-depth expertise in Kubernetes, containerized workloads, and production-scale distributed systems.

• Familiarity with CI/CD pipelines utilizing GitHub Actions, ArgoCD, Terraform, or comparable DevOps tools.

• Solid experience in monitoring, logging, tracing, and observability platforms such as Grafana, Prometheus, Datadog, or Splunk.

• Proficient in scripting and automation using Python, Bash, PowerShell, or similar programming languages.

• Strong understanding of release engineering, automated testing frameworks, QA tooling, and shift-left quality practices.

• Experience in supporting SaaS applications with uptime, scalability, and security requirements in regulated industries.

• Knowledge of HIPAA, SOC2, vulnerability management, access controls, and infrastructure security best practices.

• Familiarity with databases, APIs, networking, and troubleshooting within modern web application stacks.

• Excellent communication and cross-functional influence skills.

• Preferred qualifications include experience in healthcare technology or regulated SaaS, familiarity with FHIR-native or EMR/EHR architectures, AI-assisted SRE automation, Playwright or equivalent, and building internal SRE capabilities alongside managed services.

• Candidates must be based in the United States.

• This position is not eligible for sponsorship.


🏝️ Benefits

• A variable compensation component.

• Stock options.

• A fully remote, collaborative engineering environment.

• Direct access to executive leadership.

• An opportunity to establish the SRE function from the ground up.

• A chance to lead a blended team model.

• The opportunity to work on systems impacting clinical care.

• An AI-assisted software development environment with Claude Code.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers