
Staff Site Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in California.
• Construct, manage, and troubleshoot production Kubernetes/EKS clusters.
• Execute Kubernetes upgrades, node rollouts, and perform cluster maintenance.
• Develop and oversee AWS infrastructure, including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage solutions.
• Define and maintain infrastructure using Terraform.
• Create and manage CI/CD and deployment infrastructure.
• Diagnose production issues across Kubernetes, AWS, Linux, networking, and databases.
• Establish monitoring, alerting, and observability for critical infrastructure.
• Engage in on-call rotations and respond to production incidents.
• Identify and address infrastructure scaling and reliability challenges.
• Automate operational tasks using Python, Go, or similar programming languages.
• Assist in expanding infrastructure across new regions and deployment environments.
• Over 8 years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer, or in a comparable infrastructure role.
• Strong practical experience in operating Kubernetes.
• Experience in managing Kubernetes/EKS upgrades and production clusters.
• Solid understanding of AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.
• Production experience with Terraform or similar infrastructure-as-code tools.
• Proven experience in owning or maintaining CI/CD and deployment systems like Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.
• Background in diagnosing production infrastructure and networking issues.
• Experience in resolving significant scaling or reliability challenges.
• Verification of U.S. person status and ability to access controlled or restricted information as necessary.
• Bonus: Experience with Helm and GitOps.
• Bonus: Familiarity with Datadog or similar observability tools.
• Bonus: Experience in PostgreSQL/database operations.
• Bonus: Experience in multi-region infrastructure.
• Bonus: Experience with on-premises or disconnected deployments.
• Bonus: Experience with streaming or high-throughput distributed systems.
• Competitive base salary.
• Equity in the form of stock options.
• Comprehensive benefits package.
• Relocation assistance may be offered for eligible positions.
• Company-sponsored group health insurance plans.
• Paid vacation time.
• Sick leave.
• Holiday pay.
• 401K savings plan.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.