Senior Site Reliability Engineer

Posted 5 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the overall reliability, performance, and resilience of Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.

• Establish, assess, and maintain service level objectives (SLOs) across essential services.

• Participate in the on-call rotation and spearhead incident response efforts.

• Conduct root cause analyses, implement corrective measures, and oversee thorough infrastructure-change evaluations.

• Develop and sustain monitoring, alerting, and observability systems.

• Convert scaling needs into automated, modular Terraform infrastructure-as-code outputs.

• Execute improvements in cloud cost-efficiency and performance throughout the stack.

• Minimize operational toil and technical debt by utilizing AI tools and automation.

• Establish and uphold deployment and observability standards for engineering teams.

• Convey cloud and reliability principles to both technical and non-technical audiences.

• Ensure that infrastructure and operations adhere to security and HIPAA compliance requirements.


⛳️ Requirements

• A minimum of 4 years of practical experience managing production cloud infrastructure at scale within an SRE, DevOps, or platform engineering capacity.

• Extensive knowledge of Kubernetes and Terraform in a cloud-first setting.

• AWS experience is preferred.

• Solid background in production observability, which includes defining SLOs, developing monitoring and alerting systems, leading incident response, and carrying out blameless post-incident reviews.

• Strong foundation in software engineering, particularly in Python or Go, as it relates to infrastructure automation.

• Proven experience in enhancing cloud cost-efficiency and performance across compute, storage, and networking resources.

• Proficiency with AI tools like Claude applied within engineering and operations workflows, or a strong desire to develop this expertise swiftly.

• Must not require employer sponsorship or the transfer of an employment visa.

• Previous experience supporting AI/ML or data-intensive workloads in a production environment is advantageous.

• Experience in a security-focused or regulated environment such as HIPAA or SOC 2 is a plus.

• Familiarity with Kubernetes APIs is a plus.


🏝️ Benefits

• Equity incentive plan.

• Flexible PTO.

• Medical plan options.

• Dental plan options.

• Vision plan options.

• 401(k) with company match.

• Flexible spending accounts.

• Teladoc Health.

People also viewed

Rimutee7 hours ago

DevOps AWS – Español

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CACI International Inc7 hours ago

HCM Cloud Platform – DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$90.3k – $189.6k/year
ApplyView job
CACI International Inc7 hours ago

Senior HCM Cloud Platform – DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$105.1k – $231.1k/year
ApplyView job
Rimutee8 hours ago

DevOps, AWS – Spanish

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
OpenObserve10 hours ago

DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ShiftKey14 hours ago

Senior Site Reliability Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)PLN 22k – PLN 25k/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers