
Principal Site Reliability Engineer, Platform Engineering
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States, +2 more countries.
• Define the technical vision for GitLab Dedicated, influencing the architecture and platform strategy as we expand a growing array of isolated, single-tenant environments.
• Spearhead platform transformations encompassing resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations.
• Propel scalable, modular architecture that aligns with GitLab’s comprehensive Cells strategy while maintaining security, isolation, and compliance standards.
• Enhance service ownership and operational maturity, assisting engineering teams in building, managing, and refining the production systems they oversee.
• Identify and mitigate systemic reliability and scalability risks by utilizing production signals, incident patterns, and architectural insights.
• Develop reusable platform patterns and automation processes that minimize operational toil, enabling Dedicated to scale effectively.
• Guide intricate technical decisions across teams, carefully balancing reliability, security, cost, maintainability, and customer requirements.
• Promote engineering excellence through architectural guidance, mentorship, and influence among senior engineers and engineering leaders.
• Profound knowledge in Site Reliability, Platform, Infrastructure, or Backend Engineering, with a background in designing and managing large-scale production systems.
• Practical experience with cloud infrastructure, automation, observability, infrastructure as code, and contemporary production engineering techniques.
• Strong foundational skills in software engineering, with a history of building production systems or infrastructure tools using languages such as Go, Ruby, Python, or similar.
• Extensive expertise in distributed systems and system design, with sound judgment regarding reliability, failure isolation, scalability, and operational complexity.
• A proven record of technical leadership across various teams, establishing direction and driving complex initiatives through influence.
• Experience in leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems during substantial growth phases.
• Proven ability to lead changes that enhance how engineering teams manage and operate production systems, reinforcing reliability, operational readiness, and accountability at scale.
• Outstanding technical communication and influence skills, with the capacity to promote alignment, mentor senior engineers, and navigate complex architectural decisions.
• Benefits designed to support your health, finances, and overall well-being.
• Flexible Paid Time Off.
• Access to Team Member Resource Groups.
• Equity Compensation & Employee Stock Purchase Plan.
• Growth and Development Fund.
• Parental Leave.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.