Remotery

Senior Site Reliability Engineer – Cloud Platform

Posted Jul 11

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Ensure the reliability, availability, and performance of both production and pre-production environments.

• Oversee platform health while enhancing alerting, automation, and operational processes.

• Address production incidents, engage in root cause analysis, and execute long-term enhancements.

• Design, develop, and refine observability solutions utilizing metrics, logs, traces, and dashboards.

• Collaborate with software engineers to enhance application reliability throughout the entire development lifecycle.

• Create and maintain operational documentation, troubleshooting manuals, and runbooks.

• Streamline repetitive operational tasks through automation to boost efficiency and minimize manual involvement.

• Take part in on-call rotations and continuously refine incident response procedures.

• Advocate for reliability engineering principles, operational excellence, and ongoing improvement across engineering teams.


⛳️ Requirements

• A Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.

• Extensive experience in operating Kubernetes or other container orchestration platforms.

• Proven experience in supporting large-scale production services.

• Practical experience with AWS.

• Familiarity with Prometheus, Grafana, and ELK.

• Strong scripting abilities in Bash, Python, or Go.

• Experience in administering Linux-based production environments.

• Knowledge of Infrastructure as Code or configuration management tools like Terraform or Ansible.

• A solid grasp of networking fundamentals (TCP/IP, DNS, load balancing, routing).

• Exceptional troubleshooting, communication, and teamwork skills.

• A proactive attitude with a strong enthusiasm for automation and reliability.


🏝️ Benefits

• Long-term, full-time collaboration.

• Flexible remote working environment.

• Opportunities for professional development, including training and technical learning.

• The chance to work on cutting-edge cloud technologies used by customers globally.

• A collaborative engineering culture emphasizing knowledge sharing and continuous improvement.

• Provision of modern Apple equipment.

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers