
Senior Site Reliability Engineer
Posted Jul 20

Posted Jul 20
This is a fully remote position, open to applicants in United Kingdom.
• Take charge of enhancing the reliability, availability, and performance of essential production services across the Runware platform.
• Develop and refine our reliability methodologies, which include SLIs, SLOs, alerting mechanisms, observability, and production-readiness criteria.
• Analyze intricate production challenges within distributed systems, APIs, networking, queues, databases, and GPU-supported workloads, while participating in our engineering on-call rotation.
• Lead and engage in incident analyses and Root Cause Analyses (RCAs), transforming recurring failure patterns into sustainable engineering enhancements.
• Minimize operational toil through automation, automated remediation, and advancements in deployment safety, recovery, and system resilience.
• Collaborate closely with Engineering and DevOps teams on capacity planning, performance optimization, scaling, and architectural upgrades as the platform expands.
• Possess significant experience managing and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or a related role.
• Have a solid understanding of distributed systems and be adept at debugging across applications, databases, queues, containers, networking, and infrastructure.
• Experience in designing and managing observability frameworks utilizing metrics, logs, and distributed tracing.
• Familiar with SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
• Proficient in Kubernetes, containers, Infrastructure as Code (IaC), and automated deployment practices, with the capability to write software and automation using languages such as Python, Go, or PHP.
• Exhibit strong ownership of production issues and be comfortable participating in an engineering on-call rotation, managing problems from initial investigation to long-term solutions.
• Bonus:
• Experience in operating high-throughput or low-latency APIs and distributed systems.
• Familiarity with bare-metal infrastructure, GPU environments, or AI and machine learning workloads.
• Experience with RabbitMQ or other distributed messaging and queuing systems.
• Experience with MySQL, Redis, ClickHouse, or similar production data systems.
• Knowledge of global traffic management, load balancing, CDN platforms, and hybrid infrastructure environments.
• Experience in building automated scaling, capacity management, or self-healing systems.
• Generous paid time off – including vacation, sick days, and public holidays.
• Meaningful stock options – participate in the growth you help create.
• Remote-first environment – work from home wherever we can hire you.
• Flexible hours – manage your schedule around core collaboration times.
• Family leave – paid maternity, paternity, and caregiver leave.
• Company retreats – biannual gatherings in inspiring locations.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.