
Senior Site Reliability Engineer
Posted 21 hours ago

Posted 21 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the reliability of essential production systems from start to finish.
• Define and uphold Service Level Indicators (SLIs) and Service Level Objectives (SLOs), including dashboards, alerts, and error-budget practices.
• Enhance the quality of alerts and improve detection of anomalies and correctness.
• Lead incident response efforts, restore services, and create actionable postmortems.
• Construct secure, resilient, and cost-effective infrastructure with clear handling of failure modes.
• Conduct load testing, profiling, saturation analysis, and capacity planning.
• Increase change safety through progressive delivery, automated rollback, pre-production signals, and safe deployment methodologies.
• Manage infrastructure as code and configuration, primarily utilizing Terraform.
• Design, develop, and deploy software and tools for developers.
• Facilitate game days and chaos engineering exercises.
• Collaborate with product engineering teams on production readiness, capacity planning, failure modes, rollback procedures, runbooks, and on-call handoff processes.
• Engage in and enhance the on-call rotation.
• Apply a security perspective to engineering tasks and peer reviews.
• Act as a primary resource for complex production challenges and mentor engineers through code reviews, collaborative work, and design feedback.
• 6–10 years of experience in Site Reliability Engineering (SRE), production engineering, infrastructure, or backend engineering, preferably in cloud environments (AWS is a plus).
• Proven track record of owning a system from beginning to end.
• Practical experience in defining and managing SLIs, SLOs, and error budgets.
• Experience leading or acting as the primary responder during high-severity, customer-facing incidents.
• Knowledge of distributed systems failure modes in high-throughput, low-latency scenarios.
• Strong foundation in cloud infrastructure principles, including networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.
• Significant hands-on experience managing infrastructure through code and configuration (Terraform or similar tools).
• Proficient programming skills in Go, Python, or a similar language.
• Familiarity with observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.
• Hands-on experience with operating Redis/ElastiCache in production settings, including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies.
• Knowledge of software engineering best practices, including source control, code reviews, comprehensive testing coverage, and safe deployment.
• High level of personal ownership and autonomy, with the ability to work effectively without clearly defined requirements.
• Pragmatic approach to balancing reliability and delivery.
• Excellent written and verbal communication skills in English.
• Proficient in the AI-native use of AI tools for incident investigation, telemetry analysis, runbooks, and tooling.
• Must have authorization to work from the designated home location.
• Visa sponsorship is not available.
• 100% remote work opportunity.
• The ability to work from nearly any country, subject to local restrictions.
• Visa sponsorship is not available.
• An inclusive work environment.
• CCPA and GDPR notifications for applicable residents.
In All Media
Verity Group
Endava
CVS Health
Get handpicked remote jobs straight to your inbox weekly.