
Director, Site Reliability Engineering
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Lead and expand our SRE team of approximately 10 engineers, focusing on hiring, retention, career growth, and performance management across various time zones (US, HK, NZ).
• Establish strategic collaborations with product engineering teams — transitioning SRE from a reactive, ticket-based support model to proactive co-ownership of reliability outcomes.
• Scale our multi-tenant infrastructure to facilitate new customer onboarding and accommodate expanding patient populations.
• Manage cloud cost and FinOps practices, creating frameworks that balance cost efficiency with reliability and performance.
• Advocate for developer self-service and platform engineering. Develop self-service capabilities enabling product teams to handle routine operations independently without submitting SRE tickets. Define SLOs/SLIs for critical services and enhance alert quality to ensure every notification is actionable.
• Make certain the SRE team effectively utilizes AI tools in their workflows — employing tools like Claude Code for IaC generation, log analysis, root cause investigation, and automating repetitive tasks — at the same proficiency level as the rest of the engineering team.
• You possess over 6 years of experience managing an SRE team and more than 10 years of hands-on SRE or infrastructure engineering experience.
• You are well-versed in our core technology stack: Kubernetes, GCP (GKE, Cloud SQL, Pub/Sub, GCS), Terraform, Helm, ArgoCD, PostgreSQL, and Prometheus/Grafana.
• You have strong programming skills in Python and/or Go, and you are confident in writing and reviewing infrastructure tooling code — including utilizing AI coding tools.
• You have experience with CI/CD pipelines (GitHub Actions) and a proven history of enhancing developer tooling and automation.
• You possess sound judgment regarding build versus buy decisions — you instinctively choose the right solution, not the easiest one, and you are comfortable creating internal tools when existing options do not suffice.
• You have experience leading teams across different time zones and a track record of developing engineers into capable technical contributors.
• Financial Well-Being: Our dedication to attracting and retaining top talent starts with a competitive base salary and equity opportunities. Additionally, we provide a performance-based bonus program, 401k matching, and regular compensation reviews to acknowledge and reward outstanding contributions.
• Physical Well-Being: We prioritize the health and wellness of our employees and their families by offering comprehensive medical, dental, and vision coverage. Your health is important to us, and we invest in ensuring you have access to quality healthcare.
• Mental Well-Being: We recognize the significance of mental health in enhancing productivity and maintaining work-life balance. To support this, we offer initiatives such as No-Meeting Fridays, monthly company holidays, access to mental health resources, and a generous flexible time-off policy. Furthermore, we embrace a remote-first culture that fosters collaboration and flexibility, allowing our team members to excel from any location.
• Professional Development: Cultivating internal talent is a key focus for Clover. We provide learning programs, mentorship, professional development funding, and regular performance feedback and evaluations.
• Additional Perks: Employee Stock Purchase Plan (ESPP) offering discounted equity opportunities.
• Reimbursement for office setup expenses.
• Monthly cell phone & internet stipend.
• Remote-first culture, promoting collaboration with global teams.
• Paid parental leave for all new parents.
• And much more!
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.