Senior Reliability Engineer

Posted 15 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of reliability for Tier 1 customer journeys, encompassing checkout, telehealth visits, and prescription fulfillment.

• Establish and implement Service Level Objectives (SLOs), key performance indicators, and business-level monitoring systems.

• Enhance the detection ratio of issues identified by monitoring systems compared to those found by humans.

• Lead initiatives focused on capacity and resilience for critical events and ongoing growth.

• Design load testing scenarios and analyze past incidents with detailed second-by-second breakdowns.

• Investigate and resolve database and backend performance bottlenecks.

• Strengthen caching mechanisms, GraphQL interfaces, VPC capacity, and vendor rate limits.

• Develop and evolve FireHydrant, Datadog, and Jira into a cohesive incident-response workflow.

• Automate the creation of incident reports and Root Cause Analysis (RCA) tickets, SLO burn-rate alerts, Tier 1 alert routing, and tracking of post-mortem actions.

• Write and update runbooks along with severity classification standards.

• Build and manage AI tools and agents for generating OER reports, drafting RCAs, detecting outdated action items, identifying gaps in monitoring and runbooks, and initial incident triage.

• Troubleshoot complex cross-boundary issues that involve frontend, API, service mesh, and database layers.

• Collaborate with Security and product teams on detecting and responding to anomalies.

• Create tools and metrics for bi-weekly VP-level Operational Excellence reviews and monthly cross-engineering OERs.

• Ensure the completion of associated action items.

• Document systems, onboard new team members, and mentor engineers on SLOs, blameless post-mortems, and on-call practices.


⛳️ Requirements

• A minimum of 5 years of experience as a Software, Site Reliability Engineer (SRE), Platform, or Infrastructure Engineer.

• Proven track record of managing reliability outcomes for production systems relied upon by customers.

• Solid foundation in software engineering principles.

• Capable of solving reliability challenges through coding and tool development.

• Proficient in reading application code across different layers of the stack.

• Hands-on experience with observability and SLO engineering, including golden signals, burn-rate alerts, and journey-level monitoring.

• Experience in production environments with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora).

• Familiarity with end-to-end incident management processes, including on-call design, escalation policies, incident command, blameless post-mortems, and follow-through on action items.

• Regular practical use of AI coding and analysis tools like Claude or Cursor.

• Ability to discern when to trust AI outputs and when to validate them.

• Strong communication skills to convey risks, trade-offs, and post-incident insights to engineering teams and leadership.

• Preferred: Experience in developing AI agents or LLM-supported automation for operational tasks.

• Preferred: Load testing and performance engineering experience at significant scales, such as k6.

• Preferred: Knowledge of service mesh technologies, particularly Istio.

• Preferred: Experience in a regulated or healthcare setting.

• Preferred: Background in designing vendor and partner escalation frameworks with specified severities and response SLAs.

• Must be legally authorized to work in the U.S. without any restrictions for any employer.

• Must not require immigration sponsorship by Hims & Hers to work in the U.S.


🏝️ Benefits

• Competitive salary.

• Equity compensation.

• Unlimited paid time off (PTO).

• Company holidays.

• Quarterly mental health days.

• Medical, dental, and vision insurance benefits.

• Parental leave.

• Employee Stock Purchase Program (ESPP).

• 401k benefits with employer matching contributions.

• Offsite team retreats.

• Claude Enterprise license.

• A talent-first flexible/remote work approach.

People also viewed

Akamai Technologies7 hours ago

Senior Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BeyondTrust7 hours ago

Senior Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cencora7 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PrizePicks10 hours ago

Manager, Database Reliability Engineering

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$200k – $220k/year
ApplyView job
AbbVie13 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$109.5k – $208.5k/year
ApplyView job
onXmaps, Inc.14 hours ago

Site Reliability Engineer III

US flagColorado, +6 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $153k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers