
Senior Site Reliability Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in California.
• Develop tools to enhance SRE observability.
• Engage in the Kubernetes migration process, including VMI setup and troubleshooting.
• Quickly diagnose and address incidents and issues reported by users.
• Automate, script, and create tools for both new and existing scripts to achieve full automation of daily activities.
• Provide support for services prior to launch through system design consultation, software platform and framework development, capacity management, and launch evaluations.
• Take part in an on-call rotation to support production systems.
• Propel tools and service development to uphold and enhance service SLOs.
• Collaborate with Service Owners to ensure service reliability.
• Spearhead production enhancements, including change management, post-mortem analyses, workflow processes, and software automation.
• MS or BS in Computer Science, Engineering, or a related discipline, or equivalent experience.
• Over 8 years of experience in Site Reliability Engineering with large-scale distributed microservices in production.
• Strong expertise in Kubernetes, including complex, highly available VMI configurations.
• Proven experience in leading production improvements, change management, post-mortems, workflow processes, and software automation.
• Strong problem-solving skills and root-cause analysis capabilities.
• Familiarity with monitoring systems such as Datadog, Prometheus, Alertmanager, or similar tools.
• Experience managing multi-region cloud deployments on AWS, GCP, or Azure.
• Proficiency in designing and managing deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD.
• Production-level coding skills in Go, Python, or robust Bash scripting.
• Required hands-on experience in production on-call roles responding to and addressing high-severity infrastructure alerts and service degradations.
• Excellent communication, presentation, interpersonal, and analytical skills.
• Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms is advantageous.
• Comfortable utilizing AI daily as an SRE.
• Previous experience as an SRE or Service Engineer is a plus.
• Competitive salary package.
• Equity.
• Comprehensive benefits.
Level Data
Level Data
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.