Remotery

Senior Site Reliability Engineer

Posted Jul 30

This is a fully remote position, open to applicants in California.

📋 Description

• Ensuring the reliability, availability, and performance of production infrastructure and platform services.

• Operating and scaling Kubernetes platforms, which includes governance and support for multi-tenant workloads.

• Managing deployment workflows based on GitOps using ArgoCD and Helm.

• Driving infrastructure provisioning and change management with Terraform/Terragrunt.

• Building and supporting automation for CI/CD and deployment workflows utilizing GitHub Actions.

• Leading efforts in incident response, root cause analysis, and initiatives for post-incident improvements.

• Minimizing operational toil through scripting, tooling, and process automation.

• Advancing observability practices encompassing logs, metrics, traces, dashboards, and alerting.

• Supporting secure secrets integration, IAM-aware operations, and platform guardrails.

• Collaborating closely with application, security, and platform teams to enhance reliability and delivery outcomes.


⛳️ Requirements

• 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.

• Strong hands-on experience in operating AWS within production environments.

• In-depth expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration.

• Proven experience with Kubernetes multi-tenancy, which includes namespaces, RBAC, quotas, policies, and tenant isolation patterns.

• Experience in implementing and operating ArgoCD within a GitOps delivery framework.

• Strong hands-on experience with Helm.

• Extensive experience with Terraform/Terragrunt for infrastructure provisioning and environment management.

• Solid scripting and automation capabilities using Bash and/or Python.

• Experience in building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions.

• Strong troubleshooting abilities across Linux, containers, IAM, networking, and distributed systems.

• Experience with monitoring, alerting, and observability in production environments.

• Demonstrated ownership mindset with experience managing incidents, resolving production issues, and ensuring follow-through after outages.

• Strong collaboration and communication skills, enabling effective work across engineering, security, and platform teams.

• Bachelor’s degree in computer science, engineering, a related field, or equivalent experience.

• Proven ability to leverage AI to enhance speed and quality in daily workflows for relevant outputs.

• Strong track record of critically evaluating and verifying AI-assisted work (e.g., testing, source-checking, data validation, peer review).

• High integrity and ownership: you safeguard sensitive data, avoid excessive reliance on AI, and remain accountable for final decisions and deliverables.


🏝️ Benefits

• Equity.

• Flexibility to perform at your best.

• Opportunities for professional development.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers