
Site Reliability Engineer – Intermediate to Senior Staff, Infrastructure Platforms
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants anywhere in the world.
• Ensure that user-facing services and production systems remain reliable, scalable, and efficient.
• Develop automation and tools that minimize manual work and implement repeatable, infrastructure-as-code-driven workflows.
• Manage and troubleshoot production systems on Kubernetes, including handling deployments, rollouts, and scaling operations.
• Create and maintain infrastructure as code, deploying changes safely through CI/CD and GitOps.
• Engage in on-call duties, triaging alerts, refining runbooks, and escalating issues as necessary.
• Enhance the observability stack by utilizing metrics, logs, and SLOs to identify symptoms early rather than just reacting to outages.
• Participate in incident response and post-incident analyses, transforming insights into improvements in automation and processes.
• Document runbooks, architectural decisions, and reviews to ensure that your insights become standardized practices.
• Proven experience in maintaining the reliability of production systems, blending an operations mindset with authentic software engineering skills.
• Background in creating new infrastructure tooling and automation from scratch, rather than merely configuring existing solutions. For instance, developing Terraform modules, Kubernetes operators or controllers, or building production automation and services from the ground up.
• Proficiency in reading, debugging, and reasoning about code. Most of our teams use Go, while some utilize Ruby. You should be able to discuss code behavior, performance, and potential failure modes.
• Experience with infrastructure as code, and a solid understanding of Kubernetes and its ecosystem, commensurate with your experience level.
• Practical experience with at least one major cloud provider, such as GCP or AWS.
• Awareness of observability practices, including metrics, logging, alerting, and SLOs or SLIs, and the ability to leverage data for operational decision-making.
• Comfort in participating in on-call and incident response situations, employing a structured approach to troubleshooting under pressure.
• Excellent written communication skills and the ability to function effectively as a manager-of-one in an asynchronous, distributed setting.
• A history of utilizing automation, and increasingly AI, to streamline processes and enhance team productivity.
• Alignment with GitLab's values and a dedication to working in accordance with them.
• Benefits designed to support your health, finances, and overall well-being.
• Flexible Paid Time Off policy.
• Access to Team Member Resource Groups.
• Opportunities for Equity Compensation & Employee Stock Purchase Plan.
• A dedicated Growth and Development Fund.
• Parental Leave benefits.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.