Remotery

Site Reliability Engineer – Intermediate to Senior Staff, Infrastructure Platforms

Posted Jul 22

This is a fully remote position, open to applicants anywhere in the world.

📋 Description

• Ensure that user-facing services and production systems remain reliable, scalable, and efficient.

• Develop automation and tools that minimize manual work and implement repeatable, infrastructure-as-code-driven workflows.

• Manage and troubleshoot production systems on Kubernetes, including handling deployments, rollouts, and scaling operations.

• Create and maintain infrastructure as code, deploying changes safely through CI/CD and GitOps.

• Engage in on-call duties, triaging alerts, refining runbooks, and escalating issues as necessary.

• Enhance the observability stack by utilizing metrics, logs, and SLOs to identify symptoms early rather than just reacting to outages.

• Participate in incident response and post-incident analyses, transforming insights into improvements in automation and processes.

• Document runbooks, architectural decisions, and reviews to ensure that your insights become standardized practices.


⛳️ Requirements

• Proven experience in maintaining the reliability of production systems, blending an operations mindset with authentic software engineering skills.

• Background in creating new infrastructure tooling and automation from scratch, rather than merely configuring existing solutions. For instance, developing Terraform modules, Kubernetes operators or controllers, or building production automation and services from the ground up.

• Proficiency in reading, debugging, and reasoning about code. Most of our teams use Go, while some utilize Ruby. You should be able to discuss code behavior, performance, and potential failure modes.

• Experience with infrastructure as code, and a solid understanding of Kubernetes and its ecosystem, commensurate with your experience level.

• Practical experience with at least one major cloud provider, such as GCP or AWS.

• Awareness of observability practices, including metrics, logging, alerting, and SLOs or SLIs, and the ability to leverage data for operational decision-making.

• Comfort in participating in on-call and incident response situations, employing a structured approach to troubleshooting under pressure.

• Excellent written communication skills and the ability to function effectively as a manager-of-one in an asynchronous, distributed setting.

• A history of utilizing automation, and increasingly AI, to streamline processes and enhance team productivity.

• Alignment with GitLab's values and a dedication to working in accordance with them.


🏝️ Benefits

• Benefits designed to support your health, finances, and overall well-being.

• Flexible Paid Time Off policy.

• Access to Team Member Resource Groups.

• Opportunities for Equity Compensation & Employee Stock Purchase Plan.

• A dedicated Growth and Development Fund.

• Parental Leave benefits.

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers