Site Reliability Engineer – Infrastructure Platforms, Intermediate to Senior Staff

Posted Sep 8

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Ensure the reliability, scalability, and efficiency of user-facing services and production systems.

• Develop automation and tools that minimize manual tasks and implement repeatable workflows driven by infrastructure as code.

• Manage and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling operations.

• Create and maintain infrastructure as code, safely deploying changes through CI/CD and GitOps processes.

• Engage in on-call duties, triage alerts, enhance runbooks, and escalate issues as necessary.

• Contribute to the observability framework using metrics, logs, and service level objectives (SLOs) to identify issues early.

• Take part in incident response and post-incident evaluations, transforming insights into automation and process enhancements.

• Document runbooks, architectural decisions, and review findings.

• Work collaboratively across Infrastructure Platforms teams in an asynchronous manner, driving continuous reliability improvements at scale.


⛳️ Requirements

• This position is available exclusively for candidates located in the United Kingdom.

• Proven experience in maintaining reliable production systems, merging operational thinking with software engineering practices.

• Skilled in developing new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or bespoke production automation services.

• Proficient in reading, debugging, and analyzing code, including understanding behavior, performance, and potential failure modes.

• Experience with infrastructure as code, Kubernetes, and its associated ecosystem.

• Practical experience with at least one major cloud provider: GCP or AWS.

• Knowledge of observability practices, encompassing metrics, logging, alerting, and SLOs or SLIs.

• Comfortable with on-call responsibilities and incident response activities.

• Strong written communication skills and the ability to function independently in an asynchronous, distributed environment.

• Demonstrated ability to leverage automation and increasingly AI to alleviate manual work and enhance team workflows.

• Alignment with GitLab's values and a commitment to operating in accordance with them.


🏝️ Benefits

• Comprehensive benefits to support your health, financial needs, and overall well-being.

• Flexible Paid Time Off policy.

• Access to Team Member Resource Groups.

• Equity Compensation and Employee Stock Purchase Plan.

• Growth and Development Fund.

• Parental Leave.

People also viewed

Adapty.io1 day ago

Senior DevOps Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Oddball1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $155k/year
ApplyView job
General Dynamics Information Technology1 day ago

Senior Oracle Application DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$98.6k – $133.4k/year
ApplyView job
StarTekk1 day ago

Salesforce DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Prelim1 day ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$210k – $230k/year
ApplyView job
Akamai Technologies1 day ago

Senior Site Reliability Engineer Lead

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$146.4k – $263.6k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers