Remotery

Senior Site Reliability Engineer

Posted Jul 28

This is a fully remote position, open to applicants in United States.

📋 Description

• Establish and implement a Site Reliability Engineering (SRE) practice within a six-team organization, focusing on SLOs/SLIs, error budgets, and reliability standards that teams will genuinely adopt.

• Enhance observability through metrics, logging, distributed tracing, dashboards, and alerting to ensure that more incidents are identified through monitoring prior to external notifications.

• Minimize outage frequency by identifying systemic reliability risks and collaborating with teams to address them at their root causes.

• Decrease manual effort through automation, infrastructure-as-code, and self-service tools that teams can manage and extend independently.

• Ensure the health and usability of our observability tools, delivering documentation and training as needed.

• Conduct production readiness assessments for new services and collaborate with engineering leadership on reliability priorities and capacity planning.


⛳️ Requirements

• Over 7 years of experience in software or infrastructure engineering, with significant hands-on experience in SRE or production reliability.

• Proven success in reducing incidents and enhancing detection—the key metrics by which this role will be evaluated.

• Practical experience in defining SLOs/SLIs and utilizing error budgets to inform engineering decisions.

• Extensive expertise in observability, including metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar tools).

• Strong background in operating production systems on AWS.

• Skilled in infrastructure-as-code (e.g., Terraform) and comfortable with building automation and tools (Python, TypeScript, or similar languages).


🏝️ Benefits

• Medical, dental, and vision coverage effective from day one of employment.

• 401K plan with company match.

• Generous paid time off (PTO) in addition to paid holidays.

• Paid parental leave.

• We prioritize care: as a mental health company, we place our team's well-being at the forefront.

• Advance your career with us: refine your skills and learn new ones alongside our Learning team as Talkiatry grows.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers