Remotery

Site Reliability Engineer

Posted Jul 22

This is a fully remote position, open to applicants in Turkey.

📋 Description

• Address live incidents promptly.

• Receive automated alerts and technical escalations from Support, assessing customer impact, severity, blast radius, and the current state of the system.

• Conduct investigations using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs, and application code.

• Integrate AI into the triage and diagnosis process, validating its conclusions against tangible evidence.

• Select and implement an appropriate course of action—whether it be mitigation, rollback, repair, or a bounded fix.

• Ensure customer outcomes are restored, confirming that an alert has cleared or a dashboard has returned to a healthy state.

• Maintain visibility of ownership, uncertainties, decisions, and subsequent actions while providing Support with precise technical details for customer communication.

• Occasionally engage in customer discussions when direct technical input is beneficial.

• Facilitate an appropriate response to incidents.

• Involve the relevant product team when an incident necessitates in-depth product knowledge, significant product decisions, or comprehensive root-cause analysis.

• Escalate issues with supporting evidence, customer impact, actions already taken, and specific decisions or assistance required.

• Shield developers from routine notifications; they should generally only be alerted for genuine P0/P1 impacts or product-specific judgments that cannot be postponed.

• Create a detailed incident record and handover, ensuring that immediate mitigation, product follow-ups, and reliability-process follow-ups are directed to the appropriate owners.


⛳️ Requirements

• Agency and ownership - You take charge of ambiguous live issues, gather evidence, select a course of action, and follow through after the immediate pressure has subsided.

• Operational judgment - You can distinguish between customer impact, symptoms, and potential causes, make pragmatic decisions in uncertain situations, and recognize when intervention is no longer safe or bounded.

• Technical comfort and aptitude - You are adept at exploring unfamiliar systems through code, logs, APIs, data, infrastructure, and command-line tools, and can implement hands-on changes with a clear validation strategy.

• AI-native execution - You leverage AI for substantial technical tasks, including investigation, hypothesis generation, coding, automation, incident analysis, and workflow enhancement, all while overseeing the AI's actions and questioning its conclusions.

• Accuracy and validation discipline - You actively seek out false confidence and verify results using relevant technical and customer signals.

• Systems thinking - You identify recurring patterns and enhance the triggers, ownership, runbooks, automation, metrics, and feedback loops surrounding your work.

• Clear coordination and communication - You convey information calmly and succinctly to Support, developers, and non-technical stakeholders, making evidence, impact, uncertainty, ownership, and next steps easy to grasp.

• Curiosity and resilience - You quickly learn new products and tools, continue investigating when initial hypotheses fail, and adjust your approach based on the evidence presented.

• Previous experience managing live production systems or being part of an on-call rota is highly preferred, as it demonstrates familiarity with the challenges of incident response. However, we will also consider candidates who exhibit exceptional ownership, judgment, technical aptitude, learning speed, and performance in practical assessments.


🏝️ Benefits

• Fully remote work opportunity from anywhere in Turkey.

• Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers