
Staff Site Reliability Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of and advance the engineering-wide SRE strategy, operational model, and reliability benchmarks.
• Ensure that reliability standards are aligned with customer impact, business goals, and associated risks.
• Facilitate cross-functional alignment among Infrastructure, product, service, security, and business stakeholders.
• Enhance reliability, observability, incident response, and operational readiness across various teams.
• Establish organization-wide service ownership, meaningful SLIs and SLOs, and error budgets for critical customer journeys and services.
• Define and promote the adoption of observability standards throughout pipelines and platform components.
• Set benchmarks for dashboards, actionable alerts, runbooks, and escalation procedures.
• Lead comprehensive cross-functional reliability initiatives from start to finish.
• Elevate engineering standards related to incident management, incident command, on-call health, post-incident reviews, and recovery preparedness.
• Influence the technical direction, operational model, and growth trajectory of the SRE function.
• Participate in a 24/7 on-call rotation.
• Assist in designing a sustainable, adequately staffed, and continuously improving on-call model.
• Proven experience in designing, operating, and troubleshooting large-scale distributed systems in production settings.
• In-depth understanding of reliability engineering, observability, incident management, and production operations.
• Proven ability to translate reliability expertise into standards and practices that are embraced by others.
• Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
• Backend experience in building systems and automation to minimize operational toil, enhance safeguards, and boost operational efficiency.
• Experience in leading high-severity incidents and enhancing incident response protocols.
• Exceptional written and verbal communication skills, including the ability to create technical designs, runbooks, postmortems, and operational documentation.
• Familiarity with Python and Terraform, or similar automation and infrastructure-as-code tools.
• Experience with observability tools such as Datadog, New Relic, Grafana, or other comparable platforms.
• Experience operating production services within AWS and Kubernetes environments.
• Familiarity with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows.
• Participation in a 24/7 on-call rotation is required.
• Must be legally authorized to work in the United States.
• Must indicate if employment visa sponsorship is necessary.
• Equity package in the form of stock options for all full-time positions.
• Health insurance coverage for you and your family.
• Vision insurance coverage for you and your family.
• Dental insurance coverage for you and your family.
• Flexible vacation policy.
• Generous parental leave.
• Opportunities for career development.
• An inclusive culture.
• A collaborative environment that fosters creativity and innovative thinking.
• Fully remote work model.
• Team off-sites and in-person project kick-offs.
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
PingWind Inc. (SDVOSB)
Get handpicked remote jobs straight to your inbox weekly.