Remotery

Site Reliability Engineering Leader

Posted Jul 28

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Construct, guide, and nurture the UK SRE team while establishing operational standards, best practices, and reliability objectives.

• Guarantee the high availability, stability, security, and performance of essential platforms and services.

• Define and implement the operational strategy for our global SaaS platform, ensuring outstanding reliability, availability, and performance.

• Oversee major incident management, serving as the senior escalation point during critical production incidents.

• Establish and track reliability metrics, including SLIs, SLOs, and operational KPIs.

• Promote automation across infrastructure, deployments, monitoring, and operational workflows to enhance efficiency and minimize manual effort.

• Advocate for the integration of AI-powered operations, utilizing modern AI technologies to boost engineering productivity and operational excellence.

• Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance platform architecture, scalability, and operational preparedness.

• Direct disaster recovery planning, operational resilience initiatives, and risk management throughout the platform.

• Continuously assess emerging cloud, AI, and platform technologies to maintain PatSnap's position at the forefront of engineering excellence.


⛳️ Requirements

• Bachelor’s degree in Computer Science or a related discipline, with a minimum of 8 years of experience in DevOps, SRE, or infrastructure operations.

• Demonstrated experience in leading technical teams and managing large-scale production environments.

• Strong proficiency in cloud platforms (preferably AWS), Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code, and observability platforms.

• In-depth understanding of distributed systems, high-availability architectures, and large-scale SaaS environments.

• Experience in driving automation and initiatives for operational excellence.

• Practical experience utilizing AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar technologies to enhance engineering productivity.

• Excellent problem-solving, leadership, communication, and stakeholder management abilities.

• Proficient in English; knowledge of Mandarin is highly desired to facilitate collaboration with teams across various regions.


🏝️ Benefits

• Develop technology that drives global innovation.

• Lead AI-driven engineering initiatives.

• Manage a critical business function.

• Work with cutting-edge technologies.

• Collaborate on a global scale.

• Advance your leadership career.

People also viewed

TEKsystems20 hours ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems20 hours ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data23 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data23 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data23 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data23 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers