
Senior Reliability Engineer
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of reliability for Tier 1 customer journeys, encompassing checkout, telehealth visits, and prescription fulfillment.
• Establish and implement Service Level Objectives (SLOs), key performance indicators, and business-level monitoring systems.
• Enhance the detection ratio of issues identified by monitoring systems compared to those found by humans.
• Lead initiatives focused on capacity and resilience for critical events and ongoing growth.
• Design load testing scenarios and analyze past incidents with detailed second-by-second breakdowns.
• Investigate and resolve database and backend performance bottlenecks.
• Strengthen caching mechanisms, GraphQL interfaces, VPC capacity, and vendor rate limits.
• Develop and evolve FireHydrant, Datadog, and Jira into a cohesive incident-response workflow.
• Automate the creation of incident reports and Root Cause Analysis (RCA) tickets, SLO burn-rate alerts, Tier 1 alert routing, and tracking of post-mortem actions.
• Write and update runbooks along with severity classification standards.
• Build and manage AI tools and agents for generating OER reports, drafting RCAs, detecting outdated action items, identifying gaps in monitoring and runbooks, and initial incident triage.
• Troubleshoot complex cross-boundary issues that involve frontend, API, service mesh, and database layers.
• Collaborate with Security and product teams on detecting and responding to anomalies.
• Create tools and metrics for bi-weekly VP-level Operational Excellence reviews and monthly cross-engineering OERs.
• Ensure the completion of associated action items.
• Document systems, onboard new team members, and mentor engineers on SLOs, blameless post-mortems, and on-call practices.
• A minimum of 5 years of experience as a Software, Site Reliability Engineer (SRE), Platform, or Infrastructure Engineer.
• Proven track record of managing reliability outcomes for production systems relied upon by customers.
• Solid foundation in software engineering principles.
• Capable of solving reliability challenges through coding and tool development.
• Proficient in reading application code across different layers of the stack.
• Hands-on experience with observability and SLO engineering, including golden signals, burn-rate alerts, and journey-level monitoring.
• Experience in production environments with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora).
• Familiarity with end-to-end incident management processes, including on-call design, escalation policies, incident command, blameless post-mortems, and follow-through on action items.
• Regular practical use of AI coding and analysis tools like Claude or Cursor.
• Ability to discern when to trust AI outputs and when to validate them.
• Strong communication skills to convey risks, trade-offs, and post-incident insights to engineering teams and leadership.
• Preferred: Experience in developing AI agents or LLM-supported automation for operational tasks.
• Preferred: Load testing and performance engineering experience at significant scales, such as k6.
• Preferred: Knowledge of service mesh technologies, particularly Istio.
• Preferred: Experience in a regulated or healthcare setting.
• Preferred: Background in designing vendor and partner escalation frameworks with specified severities and response SLAs.
• Must be legally authorized to work in the U.S. without any restrictions for any employer.
• Must not require immigration sponsorship by Hims & Hers to work in the U.S.
• Competitive salary.
• Equity compensation.
• Unlimited paid time off (PTO).
• Company holidays.
• Quarterly mental health days.
• Medical, dental, and vision insurance benefits.
• Parental leave.
• Employee Stock Purchase Program (ESPP).
• 401k benefits with employer matching contributions.
• Offsite team retreats.
• Claude Enterprise license.
• A talent-first flexible/remote work approach.
Akamai Technologies
BeyondTrust
Cencora
PrizePicks
Get handpicked remote jobs straight to your inbox weekly.