
Senior Site Reliability Engineer, Cloud Platform
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United Kingdom.
• Ensure the reliability, availability, and performance of both production and pre-production environments.
• Oversee platform health and enhance alerting, automation, and operational workflows.
• Address production incidents, engage in root cause analysis, and implement sustainable improvements.
• Design, construct, and refine observability solutions utilizing metrics, logs, traces, and dashboards.
• Collaborate with software engineers to enhance application reliability throughout the development lifecycle.
• Create and maintain operational documentation, troubleshooting manuals, and runbooks.
• Automate repetitive operational tasks to boost efficiency and minimize manual intervention.
• Participate in on-call rotations while continually refining incident response processes.
• Advocate for reliability engineering principles, operational excellence, and ongoing improvement across engineering teams.
• Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.
• Significant experience managing Kubernetes or other container orchestration platforms.
• Experience in supporting large-scale production services.
• Practical experience with AWS.
• Familiarity with Prometheus, Grafana, and ELK.
• Strong scripting abilities in Bash, Python, or Go.
• Experience in administering Linux-based production environments.
• Knowledge of Infrastructure as Code or configuration management tools such as Terraform or Ansible.
• Solid grasp of networking fundamentals, including TCP/IP, DNS, load balancing, and routing.
• Exceptional troubleshooting, communication, and collaboration skills.
• A proactive attitude with a passion for automation and reliability.
• Nice to have: experience with SIP or VoIP technologies.
• Nice to have: familiarity with MySQL or PostgreSQL.
• Nice to have: experience with Redis or other NoSQL databases.
• Long-term, full-time collaboration.
• Flexible remote working environment.
• Professional development opportunities, including training and technical learning.
• Chance to work on innovative cloud technologies utilized by customers globally.
• A collaborative engineering culture emphasizing knowledge sharing and continuous improvement.
• Modern Apple equipment provided.
• An inclusive and respectful workplace.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.