
Release Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants anywhere in the world.
• Take responsibility for the reliability of Supabase's deployment and release systems, along with the control plane they operate on, adhering to established SLOs and error budgets.
• Transform pre-production into a reliable indicator by standardizing and enhancing the current fragmented, ad-hoc deployment workflows.
• Lead efforts in disaster-recovery preparedness, ensuring environments can be deployed reproducibly from scratch by disentangling undocumented secrets, clarifying configuration ownership, and resolving circular service dependencies.
• Develop and maintain health and SLO monitoring for essential user flows, utilizing synthetic testing to identify regressions before customers encounter them.
• Minimize mean-time-to-detect and mean-time-to-recover for deployment-related incidents, which constitute a significant portion of our incident workload.
• Engage in on-call duties, facilitate blameless postmortems, and transform insights into runbooks, alerts, and automation that reduce manual effort.
• Enhance deployment observability and auditability, ensuring a clear record of what was deployed, where, when, and by whom.
• Document operational procedures, including break-glass protocols, access models, and runbooks, to ensure reliability knowledge is widely available rather than limited to a few individuals.
• Establish and monitor SLAs, SLOs, error budgets, and DORA delivery metrics, providing meaningful alerts while filtering out noise.
• Guarantee that deployments fail quickly and safely when health checks indicate degradation.
• Strengthen access and break-glass workflows (e.g., scoped self-service) to empower the appropriate personnel to respond to incidents without resorting to unsafe workarounds.
• Collaborate with product engineering and platform teams to synchronize release practices with reliability and availability objectives.
• Possess 5+ years of experience in SRE, production operations, platform engineering, or release engineering.
• Have experience operating production systems at scale and handling on-call responsibilities.
• Be well-versed in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs, alongside the observability tools that support them (Prometheus, Grafana, Alertmanager, or similar).
• Have led incident responses with tools such as incident.io (or PagerDuty/Opsgenie), conducted blameless postmortems, and successfully reduced MTTD/MTTR.
• Operate confidently within AWS (multiple accounts, IAM, VPC) in a production environment.
• Be comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes.
• Automate and script to eliminate manual tasks rather than simply absorbing them.
• Communicate effectively with both infrastructure experts and product engineers.
• Excel in asynchronous, globally distributed teams.
• Feel at ease navigating ambiguity and progressively refining systems over time.
• Fully Remote
• ESOP
• Tech Allowance
• Health Benefits
• Annual Off-Sites
• Flexible Work
• Professional Development
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.