
Senior Site Reliability Engineer – Cloud Platform
Posted Jul 11

Posted Jul 11
This is a fully remote position, open to applicants in United Kingdom.
• Ensure the reliability, availability, and performance of both production and pre-production environments.
• Oversee platform health while enhancing alerting, automation, and operational processes.
• Address production incidents, engage in root cause analysis, and execute long-term enhancements.
• Design, develop, and refine observability solutions utilizing metrics, logs, traces, and dashboards.
• Collaborate with software engineers to enhance application reliability throughout the entire development lifecycle.
• Create and maintain operational documentation, troubleshooting manuals, and runbooks.
• Streamline repetitive operational tasks through automation to boost efficiency and minimize manual involvement.
• Take part in on-call rotations and continuously refine incident response procedures.
• Advocate for reliability engineering principles, operational excellence, and ongoing improvement across engineering teams.
• A Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.
• Extensive experience in operating Kubernetes or other container orchestration platforms.
• Proven experience in supporting large-scale production services.
• Practical experience with AWS.
• Familiarity with Prometheus, Grafana, and ELK.
• Strong scripting abilities in Bash, Python, or Go.
• Experience in administering Linux-based production environments.
• Knowledge of Infrastructure as Code or configuration management tools like Terraform or Ansible.
• A solid grasp of networking fundamentals (TCP/IP, DNS, load balancing, routing).
• Exceptional troubleshooting, communication, and teamwork skills.
• A proactive attitude with a strong enthusiasm for automation and reliability.
• Long-term, full-time collaboration.
• Flexible remote working environment.
• Opportunities for professional development, including training and technical learning.
• The chance to work on cutting-edge cloud technologies used by customers globally.
• A collaborative engineering culture emphasizing knowledge sharing and continuous improvement.
• Provision of modern Apple equipment.
Ontrac Solutions
CyberSheath
Ontrac Solutions
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.