Senior Site Reliability Engineer

Posted 21 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the reliability of essential production systems from start to finish.

• Define and uphold Service Level Indicators (SLIs) and Service Level Objectives (SLOs), including dashboards, alerts, and error-budget practices.

• Enhance the quality of alerts and improve detection of anomalies and correctness.

• Lead incident response efforts, restore services, and create actionable postmortems.

• Construct secure, resilient, and cost-effective infrastructure with clear handling of failure modes.

• Conduct load testing, profiling, saturation analysis, and capacity planning.

• Increase change safety through progressive delivery, automated rollback, pre-production signals, and safe deployment methodologies.

• Manage infrastructure as code and configuration, primarily utilizing Terraform.

• Design, develop, and deploy software and tools for developers.

• Facilitate game days and chaos engineering exercises.

• Collaborate with product engineering teams on production readiness, capacity planning, failure modes, rollback procedures, runbooks, and on-call handoff processes.

• Engage in and enhance the on-call rotation.

• Apply a security perspective to engineering tasks and peer reviews.

• Act as a primary resource for complex production challenges and mentor engineers through code reviews, collaborative work, and design feedback.


⛳️ Requirements

• 6–10 years of experience in Site Reliability Engineering (SRE), production engineering, infrastructure, or backend engineering, preferably in cloud environments (AWS is a plus).

• Proven track record of owning a system from beginning to end.

• Practical experience in defining and managing SLIs, SLOs, and error budgets.

• Experience leading or acting as the primary responder during high-severity, customer-facing incidents.

• Knowledge of distributed systems failure modes in high-throughput, low-latency scenarios.

• Strong foundation in cloud infrastructure principles, including networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.

• Significant hands-on experience managing infrastructure through code and configuration (Terraform or similar tools).

• Proficient programming skills in Go, Python, or a similar language.

• Familiarity with observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.

• Hands-on experience with operating Redis/ElastiCache in production settings, including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies.

• Knowledge of software engineering best practices, including source control, code reviews, comprehensive testing coverage, and safe deployment.

• High level of personal ownership and autonomy, with the ability to work effectively without clearly defined requirements.

• Pragmatic approach to balancing reliability and delivery.

• Excellent written and verbal communication skills in English.

• Proficient in the AI-native use of AI tools for incident investigation, telemetry analysis, runbooks, and tooling.

• Must have authorization to work from the designated home location.

• Visa sponsorship is not available.


🏝️ Benefits

• 100% remote work opportunity.

• The ability to work from nearly any country, subject to local restrictions.

• Visa sponsorship is not available.

• An inclusive work environment.

• CCPA and GDPR notifications for applicable residents.

People also viewed

In All Media18 hours ago

DevOps Engineer – Cloud

BR flagBrazil, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group18 hours ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Endava23 hours ago

Senior DevOps Engineer, Terraform

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CVS Health1 day ago

Staff DevSecOps Engineer, Health

US flagConnecticut, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$130.3k – $260.6k/year
ApplyView job
GoFasti1 day ago

Senior DevOps Engineer

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)$5,000 – $6,000/month
ApplyView job
Hitss Brasil1 day ago

DevOps Architect

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers