
Senior Site Reliability Engineer – Hiring Globally
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United Kingdom.
• Take ownership of service objectives. Establish SLIs and SLOs for critical services, manage error budgets, and utilize them to inform prioritization and change-rate decisions. Ensure that reliability is quantifiable rather than subjective.
• Ensure trustworthy observability. Manage alert quality throughout the process: maintain a high signal-to-noise ratio, ensure alarms effectively detect their intended incidents, and exercise the discipline to disable misleading alarms until they are accurately resolved. Create the necessary dashboards, custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) to gain clear insights into the system.
• Lead incident response efforts. Manage incidents with composure, reduce mean-time-to-recovery, and generate blameless postmortems with actionable items that are genuinely addressed. Enhance and maintain the health of the on-call rotation.
• Plan for capacity and performance. Anticipate and optimize compute resources (particularly GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Identify saturation before it impacts customers.
• Oversee business continuity and disaster recovery strategies. Manage backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and validate them through regular, rehearsed game days rather than relying on assumptions.
• Maintain the health of the edge fleet. Conduct remote diagnosis and recovery via AWS IoT, implement container auto-updates using systemd timers, and continue the CentOS 7 migration to the containerized media stack (Ubuntu 22.04).
• Eliminate operational toil. Develop software (Python, Golang, Bash) to automate routine tasks, self-correct common failures, and establish reliability as a standard practice rather than a heroic effort.
• Safeguard production changes. Enforce collaborative, reviewed change management processes; protect the system from hazardous, unilateral alterations (topology, instance-count, scaling, and configuration).
• An AWS certification is required. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is highly preferred.
• Over 10 years of experience in Site Reliability Engineering or large-scale production operations. Mastery of every area is not expected from day one, but significant depth in several areas and the ability to quickly learn the remaining ones is essential.
• Proven experience with SLO/error-budget practices. You have established SLIs and SLOs, operated within an error budget, and made concrete decisions based on it.
• Strong skills in production observability. Proficient with metrics, logs, alerting, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), and capable of building the necessary instrumentation when it is lacking.
• Demonstrated incident command experience. You have led incidents, managed an on-call rotation, and authored postmortems that have influenced system behavior.
• Experience in capacity planning and performance evaluation across compute resources, databases, and a messaging or streaming system (Kafka/MSK preferred).
• Ownership of disaster recovery processes: backups, replication, failover, and validated RPO/RTO metrics.
• Software engineering skills for automation. Proficient in writing Python, Golang, and Bash to create reliability tools, rather than just configuring existing solutions.
• Expertise in Terraform (or equivalent Infrastructure as Code) and strong Linux administration skills (shell and Linux GNU utilities), comfortable working from cloud environments to bare-metal/edge systems.
• Experience in database operations with PostgreSQL (time-series experience is a plus).
• A reliability-oriented mindset: you prioritize instrumentation over guesswork, and you document processes in runbooks.
• Nice to have:
• Experience operating GPU workloads and serving computer vision or machine learning models in production (CUDA, Deep Learning AMIs, inference scaling).
• Familiarity with Apache MSK / Kafka and operations involving streaming data (Kinesis, Kinesis Video Streams).
• Experience with AWS IoT Core at scale: device provisioning, certificates, secure tunneling.
• Managing a fleet of edge/on-premise devices (golden images, remote updates, systemd).
• Experience in operating and modernizing legacy systems (Java 8, Jetty, CentOS).
• Knowledge of chaos engineering/game-day practices and capacity modeling.
• Familiarity with Bazel in a monorepo; experience with Cloudflare, Cognito/Auth0, and API Gateway.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.