
Senior Site Reliability Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in California.
• Collaborate with teams to enhance service delivery and reliability throughout their entire lifecycle.
• Assess and track all production systems focusing on availability, latency, and overall system health.
• Investigate the root causes of errors and instability within our production cloud services, guiding teams towards improved operational excellence.
• Partner with product and platform teams to refine and enhance systems by advocating for modifications that boost reliability, resilience, and observability.
• Assist in identifying and reducing toil through innovative solutions and automation.
• This role will entail stand-by, on-call, or after-hours responsibilities.
• Demonstrated experience in designing, implementing, and managing observability systems for intricate cloud-based platforms.
• Familiarity with Configuration Management and Infrastructure as Code tools such as Terraform (preferred) or Ansible.
• Proficient in cloud platforms (preferably AWS and Azure) along with container and orchestration technologies.
• Experience with Application Performance Monitoring (APM) and observability tools including New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry, etc.
• Extensive experience working in enterprise-scale continuous delivery environments.
• Development experience with JavaScript, Node.js, or TypeScript in a Linux/Mac setting.
• Experience with sustainable incident response within a blameless culture.
• Background in Linux Systems Engineering.
• Proficiency with incident response tools such as PagerDuty, FireHydrant, Blameless, etc.
• Comfortable working autonomously and effectively within a distributed team.
• Knowledge of cloud and application security best practices.
• Strong understanding of cloud design patterns for scalability, data management, resiliency, and more.
• A passion for high-quality outcomes and an aptitude for testing.
• Insights into business metrics and Service Level Objectives (SLOs).
• Health, dental, vision, short-term disability, and life insurance.
• Paid holidays and paid time off.
• Fertility treatment benefits.
• 401(k) plan.
• Equity opportunities.
Flock Safety
Pear Tree.
GFT Technologies
Kyndryl
Get handpicked remote jobs straight to your inbox weekly.