
Lead Site Reliability Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in India.
• Engage in a 24/7 on-call rotation, swiftly addressing critical incidents to maintain high availability for the NA eCommerce platform.
• Diagnose, troubleshoot, and resolve intricate production issues as a member of incident triage teams, aiming to reduce Mean Time to Recovery (MTTR).
• Implement and enhance operational runbooks and Standard Operating Procedures (SOPs).
• Lead and partake in blameless post-mortem and Root Cause Analysis (RCA) sessions.
• Collaborate with development and platform teams to design long-term reliability solutions.
• Establish and monitor Service Level Indicators (SLIs) and Objectives (SLOs).
• Work together with Product Owners to set service levels and manage Error Budgets.
• Evaluate service health and SLO impacts during monthly release reviews.
• Utilize and optimize Dynatrace, GCP Logging, and other observability tools to assess system health and detect anomalies.
• Identify gaps in observability and implement solutions for comprehensive system visibility.
• Oversee metrics, dashboards, and alert configurations using Terraform to set up monitoring infrastructure.
• Develop notification strategies and thresholds for KPI/SLO violations.
• Create automation scripts, tools, and workflows to minimize toil.
• Design and implement self-healing mechanisms for common system failures.
• Deploy and manage AI-driven observability solutions for proactive monitoring and predictive maintenance.
• Collaborate with platform and engineering teams to address production bottlenecks and enhance processes.
• Provide data-driven reports on system health, incident trends, and SRE initiatives to leadership and stakeholders.
• Bachelor’s degree in computer science or a related field.
• At least 3+ years of professional experience in Site Reliability Engineering or DevOps.
• Extensive hands-on experience with Google Cloud Platform, particularly Cloud Run, GKE, and OpenShift.
• Advanced skills in Terraform, including creating reusable modules, managing state, and automating infrastructure provisioning.
• Experience with comprehensive observability through metrics, events, logs, and traces.
• Practical experience with Dynatrace or similar Application Performance Management (APM) tools like Datadog or New Relic, covering distributed tracing, synthetic monitoring, and code-level profiling.
• Proficiency in at least one high-level programming language: Java, Node.js, Python, or Go.
• Capability to read application code to assist with instrumentation and troubleshoot complex production issues.
• Demonstrated experience managing high-severity incidents and the incident lifecycle from triage to mitigation, resolution, and blameless post-mortem/RCA.
• Willingness to participate in a 24/7 on-call rotation.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work opportunities.
• Professional development and continuous learning opportunities.
• Generous paid time off and holiday policy.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.