
Staff Site Reliability Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants anywhere in the world.
• Collaborate closely with development teams from the beginning of projects to ensure reliability in design.
• Establish standards for production readiness and criteria for launches.
• Develop measurable Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
• Build dependable frameworks and tools across various cloud platforms.
• Oversee incident management procedures and conduct postmortem analyses.
• Foster relationships with cloud service providers to improve capabilities.
• Proficient in multiple cloud environments, including AWS, GCP, and Azure.
• Comfortable providing support for services developed in TypeScript and Ruby on Rails.
• Extensive experience as a Site Reliability Engineer (SRE) or platform engineer.
• Strong foundation in software engineering principles.
• Demonstrated technical leadership and influence without the need for formal authority.
• Capacity to drive high-impact issues to resolution.
• Ability to utilize systems thinking to recognize process and technical debt.
• Proven experience in data-driven leadership.
• Excellent verbal and written communication skills in English.
• Fully remote team.
• Collaborative work environment.
• Chance to impact reliability practices.
• Participation in on-call rotation to gain insights into live incidents.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.