
Senior Site Reliability Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in California, +7 more states.
• Participate in an on-call rotation, utilizing runbooks and playbooks to diagnose and resolve production issues.
• Design, create, and maintain observability dashboards and alerts based on Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
• Operate and enhance our Kubernetes-based computing platform.
• Collaborate across cloud networking and infrastructure (Azure/AWS) to ensure reliable and scalable systems.
• Investigate and resolve production incidents, including conducting root-cause analyses and subsequent remediation work.
• Partner with product engineering teams to evaluate architecture and infrastructure decisions prior to deployment.
• Develop and sustain automation that minimizes manual, repetitive operational tasks across the team.
• Write and maintain runbooks and documentation to disseminate on-call knowledge throughout the team.
• Assist in defining non-functional requirements—such as scalability, availability, and performance—for new system designs.
• Collaborate with engineering teams to adopt best practices in reliability and observability.
• Contribute to CI/CD pipelines and assist teams in deploying changes safely and efficiently.
• 8-10+ years of relevant hands-on experience.
• Kubernetes (essential): strong, practical understanding of Kubernetes as a system.
• SRE principles: practical experience with SLIs, SLOs, and error budgets.
• Cloud engineering & networking: solid background in AWS or Azure, including networking fundamentals (subnetting, IP addressing).
• Observability: extensive experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch).
• CI/CD: strong understanding of a CI/CD system, preferably GitHub Actions.
• Proficient programming skills with the ability to develop web applications—ideally with a strong working knowledge of .NET and ASP.NET.
• We also welcome strong backgrounds in Python (Flask, FastAPI) or Java (Spring).
• Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures).
• Strong production troubleshooting skills—capable of diagnosing issues under pressure.
• Flextime, recognition, and support for autonomous work: Flexible time off with ample learning and development opportunities to continue growing your career.
• Company-paid medical, dental, and vision (with 100% employer-paid options and 90% coverage for dependents).
• FSA and HSA, 401k match, and telehealth options including memberships to One Medical.
• Parental leave and support, up to $20k in fertility services (i.e., IUI and IVF), surrogacy, and adoption reimbursement.
• On-demand maternity support through Maven Maternity, free breast milk shipping through Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.