
Staff Site Reliability Engineer
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in United States.
• Guide the strategy for cloud architecture with an emphasis on scalability, maintainability, and cost-effectiveness.
• Lead capacity planning initiatives to proactively scale in anticipation of demand.
• Oversee the health and architecture of the Kubernetes (GKE) platform.
• Design and enhance network architecture, global routing, connectivity, and the Cloudflare edge strategy.
• Drive the implementation of Infrastructure as Code standards and develop reusable Terraform/Terragrunt modules.
• Establish and improve observability practices, including monitoring, alerting, and service level objectives (SLOs).
• Lead the planning for disaster recovery, backup design, failover strategies, and definition of recovery time objectives (RTO) and recovery point objectives (RPO).
• Architect systems that are highly available and fault-tolerant.
• Optimize cloud expenditures while balancing performance, resilience, and costs.
• Mentor senior engineers to enhance their technical decision-making skills.
• Develop tools that empower developers and establish self-service patterns.
• Represent infrastructure and reliability considerations during planning processes.
• Advocate for change management, peer reviews, staged rollouts, and go/no-go decision points.
• Lead responses to high-severity incidents and participate in an on-call rotation.
• Transform incidents into opportunities for lasting improvements in reliability.
• Possess expert-level knowledge of Google Cloud Platform (GCP) and extensive infrastructure experience.
• Demonstrate expertise in Infrastructure as Code, particularly with Terraform and Terragrunt.
• Have experience with GitLab or comparable tools and the design and operation of CI/CD pipelines.
• Exhibit expertise in Kubernetes and containerization, including Helm.
• Hold strong experience in network architecture, routing, and connectivity both within GCP and across various cloud providers.
• Have significant experience with edge security and traffic management using solutions such as Cloudflare.
• Possess deep familiarity with observability tools and monitoring strategies.
• Experience leading incident management and response initiatives is required.
• Have a proven track record at the staff-level L6, including mentoring senior engineers.
• Demonstrate experience in architecting systems that are highly available and fault-tolerant.
• Exhibit a basic ability to read, comprehend, and write code across various languages and runtimes, such as Ruby/Rails, Java, and Python.
• Current and ongoing permanent residency in the United States is mandatory.
• Experience in cloud cost allocation and control is preferred.
• Familiarity with capacity planning and infrastructure forecasting at scale is preferred.
• Knowledge of AWS and/or Azure is preferred, with a primary focus on GCP.
• Competitive benefits package.
• Opportunities for professional development.
• A collaborative workplace culture.
• Fully remote work arrangement.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.