
Site Reliability Engineer – Infrastructure Platforms, Intermediate to Senior Staff
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in United Kingdom.
• Ensure the reliability, scalability, and efficiency of user-facing services and production systems.
• Develop automation and tools that minimize manual tasks and implement repeatable workflows driven by infrastructure as code.
• Manage and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling operations.
• Create and maintain infrastructure as code, safely deploying changes through CI/CD and GitOps processes.
• Engage in on-call duties, triage alerts, enhance runbooks, and escalate issues as necessary.
• Contribute to the observability framework using metrics, logs, and service level objectives (SLOs) to identify issues early.
• Take part in incident response and post-incident evaluations, transforming insights into automation and process enhancements.
• Document runbooks, architectural decisions, and review findings.
• Work collaboratively across Infrastructure Platforms teams in an asynchronous manner, driving continuous reliability improvements at scale.
• This position is available exclusively for candidates located in the United Kingdom.
• Proven experience in maintaining reliable production systems, merging operational thinking with software engineering practices.
• Skilled in developing new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or bespoke production automation services.
• Proficient in reading, debugging, and analyzing code, including understanding behavior, performance, and potential failure modes.
• Experience with infrastructure as code, Kubernetes, and its associated ecosystem.
• Practical experience with at least one major cloud provider: GCP or AWS.
• Knowledge of observability practices, encompassing metrics, logging, alerting, and SLOs or SLIs.
• Comfortable with on-call responsibilities and incident response activities.
• Strong written communication skills and the ability to function independently in an asynchronous, distributed environment.
• Demonstrated ability to leverage automation and increasingly AI to alleviate manual work and enhance team workflows.
• Alignment with GitLab's values and a commitment to operating in accordance with them.
• Comprehensive benefits to support your health, financial needs, and overall well-being.
• Flexible Paid Time Off policy.
• Access to Team Member Resource Groups.
• Equity Compensation and Employee Stock Purchase Plan.
• Growth and Development Fund.
• Parental Leave.
Adapty.io
Oddball
General Dynamics Information Technology
StarTekk
Get handpicked remote jobs straight to your inbox weekly.