
Senior Site Reliability Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in California.
• Ensuring the reliability, availability, and performance of production infrastructure and platform services.
• Operating and scaling Kubernetes platforms, which includes governance and support for multi-tenant workloads.
• Managing deployment workflows based on GitOps using ArgoCD and Helm.
• Driving infrastructure provisioning and change management with Terraform/Terragrunt.
• Building and supporting automation for CI/CD and deployment workflows utilizing GitHub Actions.
• Leading efforts in incident response, root cause analysis, and initiatives for post-incident improvements.
• Minimizing operational toil through scripting, tooling, and process automation.
• Advancing observability practices encompassing logs, metrics, traces, dashboards, and alerting.
• Supporting secure secrets integration, IAM-aware operations, and platform guardrails.
• Collaborating closely with application, security, and platform teams to enhance reliability and delivery outcomes.
• 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
• Strong hands-on experience in operating AWS within production environments.
• In-depth expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration.
• Proven experience with Kubernetes multi-tenancy, which includes namespaces, RBAC, quotas, policies, and tenant isolation patterns.
• Experience in implementing and operating ArgoCD within a GitOps delivery framework.
• Strong hands-on experience with Helm.
• Extensive experience with Terraform/Terragrunt for infrastructure provisioning and environment management.
• Solid scripting and automation capabilities using Bash and/or Python.
• Experience in building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions.
• Strong troubleshooting abilities across Linux, containers, IAM, networking, and distributed systems.
• Experience with monitoring, alerting, and observability in production environments.
• Demonstrated ownership mindset with experience managing incidents, resolving production issues, and ensuring follow-through after outages.
• Strong collaboration and communication skills, enabling effective work across engineering, security, and platform teams.
• Bachelor’s degree in computer science, engineering, a related field, or equivalent experience.
• Proven ability to leverage AI to enhance speed and quality in daily workflows for relevant outputs.
• Strong track record of critically evaluating and verifying AI-assisted work (e.g., testing, source-checking, data validation, peer review).
• High integrity and ownership: you safeguard sensitive data, avoid excessive reliance on AI, and remain accountable for final decisions and deliverables.
• Equity.
• Flexibility to perform at your best.
• Opportunities for professional development.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.