
Senior Site Reliability Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in California.
• Take charge of the design and implementation of scalable, dependable infrastructure on GCP.
• Oversee containerized workloads utilizing Kubernetes.
• Establish monitoring, alerting, and observability frameworks in Datadog.
• Lead automation initiatives to enhance operational efficiencies.
• Create and uphold CI/CD pipelines using GitHub Actions.
• Work in partnership with data and development teams.
• A minimum of 5 years of experience in SRE, platform, or infrastructure engineering.
• Extensive hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM).
• Proficient in Docker and Kubernetes.
• In-depth knowledge of Terraform.
• Skilled in a programming language such as TypeScript, Python, Go, or a similar language.
• Capable of building actionable monitoring and alerting systems (Datadog or similar tools).
• Solid understanding of the principles of distributed systems.
• Familiarity with Git and collaborative development practices.
• Experience in incident management.
• Excellent communication abilities.
• Strong sense of ownership and accountability.
• Team-oriented mindset.
• Options for remote work.
• Paid time off.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.