Principal Site Reliability Engineer, Platform Engineering

Posted 3 days ago

This is a fully remote position, open to applicants in United States, +2 more countries.

📋 Description

• Define the technical vision for GitLab Dedicated, influencing the architecture and platform strategy as we expand a growing array of isolated, single-tenant environments.

• Spearhead platform transformations encompassing resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations.

• Propel scalable, modular architecture that aligns with GitLab’s comprehensive Cells strategy while maintaining security, isolation, and compliance standards.

• Enhance service ownership and operational maturity, assisting engineering teams in building, managing, and refining the production systems they oversee.

• Identify and mitigate systemic reliability and scalability risks by utilizing production signals, incident patterns, and architectural insights.

• Develop reusable platform patterns and automation processes that minimize operational toil, enabling Dedicated to scale effectively.

• Guide intricate technical decisions across teams, carefully balancing reliability, security, cost, maintainability, and customer requirements.

• Promote engineering excellence through architectural guidance, mentorship, and influence among senior engineers and engineering leaders.


⛳️ Requirements

• Profound knowledge in Site Reliability, Platform, Infrastructure, or Backend Engineering, with a background in designing and managing large-scale production systems.

• Practical experience with cloud infrastructure, automation, observability, infrastructure as code, and contemporary production engineering techniques.

• Strong foundational skills in software engineering, with a history of building production systems or infrastructure tools using languages such as Go, Ruby, Python, or similar.

• Extensive expertise in distributed systems and system design, with sound judgment regarding reliability, failure isolation, scalability, and operational complexity.

• A proven record of technical leadership across various teams, establishing direction and driving complex initiatives through influence.

• Experience in leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems during substantial growth phases.

• Proven ability to lead changes that enhance how engineering teams manage and operate production systems, reinforcing reliability, operational readiness, and accountability at scale.

• Outstanding technical communication and influence skills, with the capacity to promote alignment, mentor senior engineers, and navigate complex architectural decisions.


🏝️ Benefits

• Benefits designed to support your health, finances, and overall well-being.

• Flexible Paid Time Off.

• Access to Team Member Resource Groups.

• Equity Compensation & Employee Stock Purchase Plan.

• Growth and Development Fund.

• Parental Leave.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers