Remotery

Principal Site Reliability Engineer

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop and implement the long-term vision for the Kubernetes platform across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise settings

• Drive architectural decisions regarding cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost efficiency

• Lead large-scale platform projects across various engineering teams, defining technical direction, engineering standards, and measurable results

• Establish and enhance reliability practices utilizing service level objectives, service level indicators, and error budget frameworks

• Create automation-first infrastructure through Infrastructure as Code, GitOps workflows, self-healing systems, and internal platform tools

• Advocate for the responsible integration of AI-driven engineering capabilities

• Manage critical platform incidents, promote post-incident improvements, and enhance platform resilience

• Mentor senior engineers, shape technical strategy, and uplift engineering excellence through architecture reviews, coaching, and technical leadership


⛳️ Requirements

• Bachelor's Degree in Computer Science or a relevant technical discipline

• Minimum of 8 years of experience in designing, operating, and scaling distributed cloud and on-premise infrastructures

• At least 3 years of experience at the Staff, Principal, or equivalent technical leadership level

• Demonstrated experience leading large-scale infrastructure or platform projects that require cross-functional collaboration and long-term technical ownership

• In-depth knowledge of Kubernetes, including cluster architecture, networking, storage, security, operators, lifecycle management, and large-scale production operations

• Extensive experience in building and managing production infrastructures in AWS and Google Cloud Platform using Infrastructure as Code technologies such as Terraform, Pulumi, or similar tools

• Strong background in software development using Go, Python, or both

• Proficiency in GitOps, continuous integration and continuous delivery, observability, distributed systems, Linux, and reliability engineering principles

• Experience integrating AI-based tools into engineering processes

• Outstanding communication and leadership abilities, with a proven track record of mentoring engineers, influencing technical strategy, and promoting engineering excellence

• Experience in regulated industries, hybrid cloud environments, contributing to open-source projects, or holding cloud certifications is preferred

• May need to obtain a gaming license from the relevant state agency as a condition of employment


🏝️ Benefits

• Bonus

• Equity

• Benefits as applicable

• Support through the gaming license process if relevant to the position

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers