Senior Site Reliability Engineer, Observability

Posted 1 day ago

This is a fully remote position, open to applicants in United Kingdom, +1 more country.

📋 Description

• Define and enhance the observability strategy, target architecture, and platform standards.

• Establish policies for instrumentation, telemetry destinations, retention, cost, and governance.

• Design Grafana dashboards and alerts related to services and customer experiences.

• Set SLOs, SLIs, and error budgets for services that directly affect customers.

• Streamline alerting to ensure every page is actionable, assigned, and documented with a runbook.

• Lead the transition to a modern incident response and on-call platform integrated with Microsoft Teams.

• Design rotation, escalation, and post-incident review processes.

• Advocate for observability-first practices throughout the engineering team.

• Develop a telemetry governance policy that addresses Log Analytics tables, retention, sampling, and cardinality.

• Select, implement, and transition to new incident response and on-call tools before the current tool reaches its end-of-life.

• Define SLOs and SLIs alongside dashboards and burn-rate alerts.

• Design and deploy OpenTelemetry instrumentation across .NET services and collector pipelines.

• Manage Azure Monitor, Log Analytics, and Application Insights configuration as code, prioritizing Bicep.

• Oversee Grafana data sources, dashboards-as-code, alert rules, and team conventions.

• Operate incident response and on-call tools, managing routing, escalation, integrations, and Teams workflows.

• Contribute to observability standards for Kubernetes and Prometheus.

• Lead incident response efforts, conduct blameless post-incident reviews, and translate findings into engineering tasks.

• Monitor and decrease MTTR, alert volume, incident frequency, and telemetry costs per service.

• Assess observability tools and AI-assisted operations capabilities, providing recommendations.

• Maintain standards, runbooks, decision records, and rationales for thresholds.

• Participate in a light-touch shared on-call rotation.


⛳️ Requirements

• 6+ years of engineering experience, with at least 3 years in a reliability, observability, or production-operations role responsible for outcomes.

• Experience in defining SLOs, SLIs, and error budgets for real services.

• Proven track record of reducing alert noise and optimizing alerting systems.

• Hands-on experience with OpenTelemetry, including service instrumentation, collector operations, and making sampling and cardinality decisions.

• Experience managing telemetry costs and retention governance.

• Extensive knowledge of Azure Monitor, Log Analytics, and Application Insights, including KQL proficiency.

• Strong skills in Grafana, including dashboard creation, alerting, and managing these as code.

• Familiarity with infrastructure-as-code; Bicep is preferred, with transferable strong Terraform experience.

• Excellent written and spoken communication skills at C1 level or above, or native-level business English.

• Experience with selecting, implementing, or migrating incident-management and on-call platforms is a plus.

• Knowledge of Prometheus-based monitoring and alerting, including exporters, recording rules, and cardinality management is a bonus.

• Experience with Kubernetes in production, especially AKS, is advantageous.

• Background in software development, ideally with .NET, is preferred.

• Familiarity with AI-assisted operations and coding tools is a plus.

• Azure certifications AZ-104, AZ-400, AZ-305, or CKA are advantageous.

• Experience in a small or scale-up environment where you were responsible for the observability function is a plus.


🏝️ Benefits

• Guaranteed £2,000 annual pay increase, separate from merit or promotion raises.

• 26 days of holiday plus public holidays, with the option to swap public holidays.

• A day off for your birthday and work anniversary each year.

• Up to 4 weeks of remote work from any location worldwide each year.

• Genuine flexibility with fully remote work options.

• Generous parental leave policy.

• Comprehensive medical scheme including remote GP services, dental, optical, and diagnostics for UK employees.

• Access to confidential counseling, therapy, and coaching through Support Room.

• Help@Hand employee assistance program.

• Life insurance coverage of 5x salary through Unum.

• Up to 10% pension matching for UK employees.

• Annual budget for learning and development per function.

• Occasional in-person team events throughout the year.

• Asynchronous-friendly work environment.

• Shared, light on-call rotation.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers