
Principal Site Reliability Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Develop and implement the long-term vision for the Kubernetes platform across Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, and on-premise settings
• Drive architectural decisions regarding cluster lifecycle management, networking, identity and access management, observability, autoscaling, capacity planning, and cost efficiency
• Lead large-scale platform projects across various engineering teams, defining technical direction, engineering standards, and measurable results
• Establish and enhance reliability practices utilizing service level objectives, service level indicators, and error budget frameworks
• Create automation-first infrastructure through Infrastructure as Code, GitOps workflows, self-healing systems, and internal platform tools
• Advocate for the responsible integration of AI-driven engineering capabilities
• Manage critical platform incidents, promote post-incident improvements, and enhance platform resilience
• Mentor senior engineers, shape technical strategy, and uplift engineering excellence through architecture reviews, coaching, and technical leadership
• Bachelor's Degree in Computer Science or a relevant technical discipline
• Minimum of 8 years of experience in designing, operating, and scaling distributed cloud and on-premise infrastructures
• At least 3 years of experience at the Staff, Principal, or equivalent technical leadership level
• Demonstrated experience leading large-scale infrastructure or platform projects that require cross-functional collaboration and long-term technical ownership
• In-depth knowledge of Kubernetes, including cluster architecture, networking, storage, security, operators, lifecycle management, and large-scale production operations
• Extensive experience in building and managing production infrastructures in AWS and Google Cloud Platform using Infrastructure as Code technologies such as Terraform, Pulumi, or similar tools
• Strong background in software development using Go, Python, or both
• Proficiency in GitOps, continuous integration and continuous delivery, observability, distributed systems, Linux, and reliability engineering principles
• Experience integrating AI-based tools into engineering processes
• Outstanding communication and leadership abilities, with a proven track record of mentoring engineers, influencing technical strategy, and promoting engineering excellence
• Experience in regulated industries, hybrid cloud environments, contributing to open-source projects, or holding cloud certifications is preferred
• May need to obtain a gaming license from the relevant state agency as a condition of employment
• Bonus
• Equity
• Benefits as applicable
• Support through the gaming license process if relevant to the position
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.