Remotery

Senior Reliability Engineer

Posted Jul 23

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead high-impact Root Cause Analysis (RCA) investigations for intricate production incidents involving software, infrastructure, industrial controls, and execution of production SOPs.

• Facilitate structured, blameless RCA reviews in accordance with ITIL Problem Management principles — focusing on fact-based analysis, clear accountability, and prompt resolutions.

• Act as the customer-facing technical lead for RCA discussions, providing updates and formal deliverables within established SLA timelines.

• Examine logs, telemetry, and incident trends across various issues to identify recurring failure patterns, and influence teams throughout the organization to resolve them.

• Present findings, risks, and recommendations to senior internal leadership and customer stakeholders, supported by data you manage from start to finish.

• Lead Continuous Service Improvement initiatives aimed at significantly reducing repeat incidents and investigation efforts through automation, tooling, and reporting.

• Mentor colleagues and enhance RCA quality standards as a senior individual contributor.


⛳️ Requirements

• A minimum of 8 years of experience in supporting complex, business-critical production environments, with a professional focus on reliability.

• At least 5 years of experience leading technical RCA, post-incident reviews, or ITIL-aligned Problem Management across software, infrastructure, systems, or industrial technology sectors.

• Strong practical skills in troubleshooting and data analysis within large-scale distributed systems, on-premise infrastructure, custom software, logs, telemetry, and incident datasets.

• A proven history of regular interaction with customers — you should be able to specify who you engaged with, at what level, and how frequently — while building trust with executives, both technical and non-technical.

• Demonstrated capability to manage multiple high-priority investigations simultaneously while influencing cross-functional teams, without direct authority, to ensure timely completion of actions.

• Bachelor’s degree in a technical discipline, or equivalent practical experience.


🏝️ Benefits

• Medical

• Dental

• Vision

• Disability

• 401K

• PTO

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers