
Senior Site Reliability Engineer
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in California, +2 more states.
• Establish the strategy for Service Level Objectives (SLOs) and Error Budgets.
• Create intricate telemetry pipelines to achieve comprehensive full-stack observability.
• Design and oversee the enterprise standards for Infrastructure as Code (IaC).
• Build custom tools to streamline complex recovery processes and facilitate system scaling.
• Serve as the Incident Commander during significant system outages, orchestrating the technical response and guiding the Root Cause Analysis (RCA) process.
• Spearhead the integration of security-as-code within DevSecOps pipelines, ensuring adherence to RMF and NIST 800-53 standards.
• Offer technical expertise and mentorship to Mid-Level SREs and developers, promoting a culture of reliability throughout the organization.
• Over 7 years of experience in SRE or DevOps, with substantial experience in distributed systems.
• Proficiency in Go, Python, or Java, along with advanced understanding of Linux internals.
• Extensive experience in managing production Kubernetes environments and intricate cloud architectures.
• Demonstrated success in defining and achieving SLOs for high-availability systems.
• Familiarity with navigating government Risk Management Framework (RMF) processes.
• Bachelor’s or Master’s degree in Computer Science or Engineering.
• Certifications: CKA (Certified Kubernetes Administrator) and preferred industry observability certification.
• Equal opportunity employer
• Inclusive work environment
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.