Remotery

Senior Site Reliability Engineer

Posted Jul 18

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Collaborate with diverse engineering teams at ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse.

• Establish and oversee service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.

• Ensure that all infrastructure components within ClickHouse Cloud (including Dataplane, Control Plane, and ClickHouse Core) have monitoring and alerting systems in place to facilitate timely detection and resolution of incidents.

• Enhance and refine incident response procedures and post-mortem analyses for any outages in ClickHouse Cloud, including collaboration with the support team to communicate with affected customers.

• Continuously work to improve the reliability and performance of our ClickHouse services.

• Plan, enable, and drive Chaos initiatives across Engineering teams based on internal priorities.

• Manage on-call processes to address performance and reliability issues, and establish best practices for coordinating escalations to resolve problems and minimize downtime.


⛳️ Requirements

• Bachelor’s or Master’s degree in Computer Science or a related field.

• A minimum of 8 years of experience in Site Reliability Engineering or a related discipline.

• Prior experience utilizing ClickHouse in a production environment.

• Proficient in Go and/or Python.

• Strong understanding of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.

• Excellent comprehension of distributed databases and SQL, particularly ClickHouse, is a significant advantage.

• Practical experience with container orchestration tools like Kubernetes or Docker Swarm.

• Extensive experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.

• Demonstrated problem-solving abilities and solid production debugging skills.

• A passion for efficiency, availability, scalability, and data governance.

• Ability to thrive in a fast-paced environment, viewing yourself as a partner with the business in the shared goal of progressing the organization.

• A strong sense of responsibility, ownership, and accountability.

• Excellent communication and interpersonal skills.


🏝️ Benefits

• Flexible work environment - ClickHouse is a globally distributed company and remote-friendly, currently operating in over 20 countries.

• Healthcare - Employer contributions towards your healthcare.

• Equity in the company - Every new team member who joins our company receives stock options.

• Time off - Flexible time off in the US, with generous entitlements in other countries.

• A $500 home office setup for remote employees.

• Global Gatherings – We believe in the power of in-person connection and provide opportunities to engage with colleagues at company-wide offsites.

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers