Remotery

Senior Site Reliability Engineer

Posted Jul 24

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of reliability and observability throughout the organization, encompassing SLAs/SLOs, instrumentation, and on-call health.

• Design, implement, and sustain Kubernetes infrastructure utilizing Helm, EKS/ECS, and core AWS services.

• Oversee RDS/Aurora Postgres, networking, and scaling operations.

• Construct and maintain CI/CD pipelines using GitHub Actions and Argo/Helm.

• Promote the adoption of Infrastructure as Code standards with Terraform and increasingly with Crossplane.

• Collaborate with over 16 engineering teams to implement standards, tools, and processes.

• Engage in the on-call rotation and lead incident response and root-cause analysis for inter-team issues.

• Utilize AI tools daily and assist less AI-proficient teammates in adopting similar methodologies.

• Represent the Site Reliability Engineering (SRE) perspective in planning discussions with engineering leadership.

• Collaborate with a Staff-level architect and engineering leadership on significant decisions while overseeing execution and cross-team implementation.


⛳️ Requirements

• 8–10 years of experience in a Senior SRE, DevOps, or Infrastructure Engineer role.

• Background in smaller to mid-sized, high-growth SaaS companies from Series A/B through C/D.

• Extensive hands-on experience with Kubernetes, EKS or ECS, GitHub Actions, Terraform, CI/CD pipelines, and Helm/Argo.

• Database expertise with Postgres and DocumentDB (Mongo) within RDS/Aurora.

• Proven history of cross-functional collaboration, timeline management, expectation setting, and working effectively with engineers from other teams.

• Experience using AI in coding, debugging, and design workflows.

• Crossplane experience is preferred.

• Familiarity with Datadog is preferred.

• Demonstrated proactivity and a bias toward action.

• Comfortable driving change across teams with diverse priorities and resistance.

• Ability to articulate technical decisions in detail, including the rationale and methodology.


🏝️ Benefits

• Comprehensive health, dental, and vision coverage for you and your family.

• Life insurance provided.

• Mental wellness support.

• Assistance for fertility and family growth.

• Flex Time Off in addition to company-paid holidays.

• Policies for paid family leave, medical leave, and bereavement leave.

• Retirement savings plans available.

• Allowance for customizing your work and technology setup at home.

• Annual stipend for professional development.

• Equity options included.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers