
Senior Site Reliability Engineer, Observability
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United Kingdom, +1 more country.
• Define and enhance the observability strategy, target architecture, and platform standards.
• Establish policies for instrumentation, telemetry destinations, retention, cost, and governance.
• Design Grafana dashboards and alerts related to services and customer experiences.
• Set SLOs, SLIs, and error budgets for services that directly affect customers.
• Streamline alerting to ensure every page is actionable, assigned, and documented with a runbook.
• Lead the transition to a modern incident response and on-call platform integrated with Microsoft Teams.
• Design rotation, escalation, and post-incident review processes.
• Advocate for observability-first practices throughout the engineering team.
• Develop a telemetry governance policy that addresses Log Analytics tables, retention, sampling, and cardinality.
• Select, implement, and transition to new incident response and on-call tools before the current tool reaches its end-of-life.
• Define SLOs and SLIs alongside dashboards and burn-rate alerts.
• Design and deploy OpenTelemetry instrumentation across .NET services and collector pipelines.
• Manage Azure Monitor, Log Analytics, and Application Insights configuration as code, prioritizing Bicep.
• Oversee Grafana data sources, dashboards-as-code, alert rules, and team conventions.
• Operate incident response and on-call tools, managing routing, escalation, integrations, and Teams workflows.
• Contribute to observability standards for Kubernetes and Prometheus.
• Lead incident response efforts, conduct blameless post-incident reviews, and translate findings into engineering tasks.
• Monitor and decrease MTTR, alert volume, incident frequency, and telemetry costs per service.
• Assess observability tools and AI-assisted operations capabilities, providing recommendations.
• Maintain standards, runbooks, decision records, and rationales for thresholds.
• Participate in a light-touch shared on-call rotation.
• 6+ years of engineering experience, with at least 3 years in a reliability, observability, or production-operations role responsible for outcomes.
• Experience in defining SLOs, SLIs, and error budgets for real services.
• Proven track record of reducing alert noise and optimizing alerting systems.
• Hands-on experience with OpenTelemetry, including service instrumentation, collector operations, and making sampling and cardinality decisions.
• Experience managing telemetry costs and retention governance.
• Extensive knowledge of Azure Monitor, Log Analytics, and Application Insights, including KQL proficiency.
• Strong skills in Grafana, including dashboard creation, alerting, and managing these as code.
• Familiarity with infrastructure-as-code; Bicep is preferred, with transferable strong Terraform experience.
• Excellent written and spoken communication skills at C1 level or above, or native-level business English.
• Experience with selecting, implementing, or migrating incident-management and on-call platforms is a plus.
• Knowledge of Prometheus-based monitoring and alerting, including exporters, recording rules, and cardinality management is a bonus.
• Experience with Kubernetes in production, especially AKS, is advantageous.
• Background in software development, ideally with .NET, is preferred.
• Familiarity with AI-assisted operations and coding tools is a plus.
• Azure certifications AZ-104, AZ-400, AZ-305, or CKA are advantageous.
• Experience in a small or scale-up environment where you were responsible for the observability function is a plus.
• Guaranteed £2,000 annual pay increase, separate from merit or promotion raises.
• 26 days of holiday plus public holidays, with the option to swap public holidays.
• A day off for your birthday and work anniversary each year.
• Up to 4 weeks of remote work from any location worldwide each year.
• Genuine flexibility with fully remote work options.
• Generous parental leave policy.
• Comprehensive medical scheme including remote GP services, dental, optical, and diagnostics for UK employees.
• Access to confidential counseling, therapy, and coaching through Support Room.
• Help@Hand employee assistance program.
• Life insurance coverage of 5x salary through Unum.
• Up to 10% pension matching for UK employees.
• Annual budget for learning and development per function.
• Occasional in-person team events throughout the year.
• Asynchronous-friendly work environment.
• Shared, light on-call rotation.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.