
Director of SRE
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in United States.
• Take ownership of and implement the SRE strategy along with a multi-quarter roadmap focusing on reliability, observability, incident management, QA maturity, and release engineering.
• Define, assess, and enhance SLAs, SLOs, error budgets, uptime, performance, and operational health metrics.
• Direct efforts towards production reliability, which includes monitoring, alerting, on-call operations, incident response, root cause analysis, and reducing MTTR.
• Establish standards for release readiness, safety controls for deployment, and quality gates.
• Oversee external SRE vendors and partners, managing service delivery, SLA governance, escalations, performance evaluations, and compliance expectations.
• Lead the QA engineering strategy with a focus on automation, prevention of regression, test coverage, and minimizing defects in production.
• Collaborate with Security and Engineering leaders on cloud infrastructure, CI/CD pipelines, operational tooling, HIPAA, SOC2, and internal security protocols.
• Supervise Azure AKS, Kubernetes, GitOps workflows, CI/CD pipelines, GitHub Actions, secrets management, access controls, and audit readiness.
• Promote the maturity of observability through tools such as Grafana, Prometheus, logging platforms, tracing tools, and automated alerting frameworks.
• Work alongside Product, Platform, and Engineering teams to integrate reliability and quality practices throughout the software development lifecycle.
• Build, mentor, and expand the SRE and QA teams.
• Propel AI-enabled automation and intelligent tools to reduce manual toil and enhance operational excellence.
• A minimum of 12 years of experience in SRE, infrastructure, or platform engineering.
• At least 5 years of experience in engineering leadership roles.
• Demonstrated ownership of site reliability for complex, multi-tenant SaaS platforms with demanding availability standards.
• Proven experience in defining SLA and SLO frameworks, error budgets, and incident management processes at scale.
• Experience managing infrastructure or SRE service vendors, including SLA governance and performance management.
• Background in leading QA or quality engineering functions, test automation maturity, and ownership of release gates.
• Strong hands-on experience with Microsoft Azure, ideally including AKS, networking, storage, IAM, and security services.
• In-depth expertise in Kubernetes, containerized workloads, and production-scale distributed systems.
• Familiarity with CI/CD pipelines utilizing GitHub Actions, ArgoCD, Terraform, or comparable DevOps tools.
• Solid experience in monitoring, logging, tracing, and observability platforms such as Grafana, Prometheus, Datadog, or Splunk.
• Proficient in scripting and automation using Python, Bash, PowerShell, or similar programming languages.
• Strong understanding of release engineering, automated testing frameworks, QA tooling, and shift-left quality practices.
• Experience in supporting SaaS applications with uptime, scalability, and security requirements in regulated industries.
• Knowledge of HIPAA, SOC2, vulnerability management, access controls, and infrastructure security best practices.
• Familiarity with databases, APIs, networking, and troubleshooting within modern web application stacks.
• Excellent communication and cross-functional influence skills.
• Preferred qualifications include experience in healthcare technology or regulated SaaS, familiarity with FHIR-native or EMR/EHR architectures, AI-assisted SRE automation, Playwright or equivalent, and building internal SRE capabilities alongside managed services.
• Candidates must be based in the United States.
• This position is not eligible for sponsorship.
• A variable compensation component.
• Stock options.
• A fully remote, collaborative engineering environment.
• Direct access to executive leadership.
• An opportunity to establish the SRE function from the ground up.
• A chance to lead a blended team model.
• The opportunity to work on systems impacting clinical care.
• An AI-assisted software development environment with Claude Code.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.