
Site Reliability Engineer – ClickHouse
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in United Kingdom.
• Overseeing substantial fleets of EC2-based virtual machines, disks, and networking for data-heavy applications.
• Enhancing operational tools related to deployments, schema modifications, backups, restorations, and incident management.
• Collaborating closely with ClickHouse engineers to translate database-level requirements into infrastructure-level solutions.
• Decreasing operational burdens by pinpointing recurring challenges and resolving them through code and self-healing automation.
• Engaging in on-call duties and incident management, with a strong emphasis on reducing the frequency of incidents over time.
• You will have the opportunity to design and automate processes rather than merely reacting to alerts.
• Previous experience with ClickHouse or other OLAP database systems.
• Extensive experience managing production infrastructure on AWS.
• Practical experience with VM-based systems (EC2), rather than solely managed PaaS.
• Experience in automating infrastructure utilizing tools like Terraform, Ansible, or similar.
• Strong understanding of Linux systems, including disk, memory, networking, and failure modes.
• Experience in supporting stateful systems such as databases, queues, and storage systems.
• Capability to troubleshoot and analyze performance and reliability issues in a production environment.
• Comfort in taking ownership of systems end-to-end, including on-call obligations.
• Transparency: Everyone has access to our roadmap, compensation practices, strategies, and operational methods through our public company handbook. Internally, we disclose revenue, board meeting notes and slides, as well as fundraising plans so that everyone is equipped with the context necessary for sound decision-making.
• Autonomy: We empower our team members to decide their next tasks based on what will most significantly impact our customers and what they find stimulating and engaging. Engineers lead product teams and make product-related decisions. Teams are adaptable and can be restructured as needed.
• Shipping fast: We prioritize rapid product development. The company is structured around small, autonomous, and highly efficient teams of talented engineers who can deliver products faster than larger companies because they maintain end-to-end ownership of their projects.
• Time for building: Our operations do not revolve around meetings. As a fully remote company, we prefer asynchronous communication – PRs > Issues > Slack. Tuesdays and Thursdays are designated as meeting-free days, allowing us to focus on building rather than achieving perfect coordination. This will be the most productive job you’ve ever had.
• Ambition: We aim to tackle significant challenges. We firmly believe that striving for the highest possible outcomes, even if we sometimes fall short, is preferable to not attempting at all. We maintain an optimistic outlook regarding our potential and our capacity to achieve it.
• Being weird: Embracing uniqueness involves reimagining an already exceptional website for the fifth time, releasing every product related to customer data, and creating an objectively unnecessary developer tool with questionable shareholder value. Engaging in unconventional activities serves as a competitive advantage and is enjoyable.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.