
Site Reliability Engineer
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in Turkey.
• Address live incidents promptly.
• Receive automated alerts and technical escalations from Support, assessing customer impact, severity, blast radius, and the current state of the system.
• Conduct investigations using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs, and application code.
• Integrate AI into the triage and diagnosis process, validating its conclusions against tangible evidence.
• Select and implement an appropriate course of action—whether it be mitigation, rollback, repair, or a bounded fix.
• Ensure customer outcomes are restored, confirming that an alert has cleared or a dashboard has returned to a healthy state.
• Maintain visibility of ownership, uncertainties, decisions, and subsequent actions while providing Support with precise technical details for customer communication.
• Occasionally engage in customer discussions when direct technical input is beneficial.
• Facilitate an appropriate response to incidents.
• Involve the relevant product team when an incident necessitates in-depth product knowledge, significant product decisions, or comprehensive root-cause analysis.
• Escalate issues with supporting evidence, customer impact, actions already taken, and specific decisions or assistance required.
• Shield developers from routine notifications; they should generally only be alerted for genuine P0/P1 impacts or product-specific judgments that cannot be postponed.
• Create a detailed incident record and handover, ensuring that immediate mitigation, product follow-ups, and reliability-process follow-ups are directed to the appropriate owners.
• Agency and ownership - You take charge of ambiguous live issues, gather evidence, select a course of action, and follow through after the immediate pressure has subsided.
• Operational judgment - You can distinguish between customer impact, symptoms, and potential causes, make pragmatic decisions in uncertain situations, and recognize when intervention is no longer safe or bounded.
• Technical comfort and aptitude - You are adept at exploring unfamiliar systems through code, logs, APIs, data, infrastructure, and command-line tools, and can implement hands-on changes with a clear validation strategy.
• AI-native execution - You leverage AI for substantial technical tasks, including investigation, hypothesis generation, coding, automation, incident analysis, and workflow enhancement, all while overseeing the AI's actions and questioning its conclusions.
• Accuracy and validation discipline - You actively seek out false confidence and verify results using relevant technical and customer signals.
• Systems thinking - You identify recurring patterns and enhance the triggers, ownership, runbooks, automation, metrics, and feedback loops surrounding your work.
• Clear coordination and communication - You convey information calmly and succinctly to Support, developers, and non-technical stakeholders, making evidence, impact, uncertainty, ownership, and next steps easy to grasp.
• Curiosity and resilience - You quickly learn new products and tools, continue investigating when initial hypotheses fail, and adjust your approach based on the evidence presented.
• Previous experience managing live production systems or being part of an on-call rota is highly preferred, as it demonstrates familiarity with the challenges of incident response. However, we will also consider candidates who exhibit exceptional ownership, judgment, technical aptitude, learning speed, and performance in practical assessments.
• Fully remote work opportunity from anywhere in Turkey.
• Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.