Remotery

Site Reliability Engineer

atTinybirdRemoteES flagSpainFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior€58k – €97k/year

Posted 7 hours ago

This is a fully remote position, open to applicants in Spain.

📋 Description

• As a member of the Platform team, your role will concentrate on the systems that ensure Tinybird remains reliable, efficient, observable, and scalable.

• Enhancing high availability and elasticity so that the system can scale automatically and efficiently.

• Improving our observability capabilities, ranging from low-level resource usage to high-level service metrics.

• Advancing disaster recovery with improved tools, incident discovery, and enhanced on-call experiences.

• Managing Kubernetes lifecycle tasks, overseeing cluster infrastructure, autoscaling, and ensuring safe deployments.

• Gaining an understanding of how ClickHouse functions internally and maximizing its performance.

• Identifying performance bottlenecks and enhancing efficiency across storage, networking, and computing.

• Minimizing operational burdens by transforming manual or fragile processes into repeatable, well-managed systems.

• Assisting with incident prevention, conducting operational reviews, and managing follow-up tasks after reliability concerns.

• Fortifying CI/CD foundations to empower teams to build and deploy changes with greater confidence.


⛳️ Requirements

• You possess strong experience in designing, building, and operating distributed cloud architectures and large-scale web-based production systems.

• You have extensive knowledge of Kubernetes, which is crucial for this position.

• You are proficient in AWS and GCP.

• Proficiency in coding is required.

• You are comfortable working close to production: debugging incidents, understanding system behavior, improving observability, and enhancing service reliability.

• You think systemically and pay close attention to edge cases, failure modes, and specific implementation details.

• You value performance, reliability, cost efficiency, and operational simplicity.

• You take ownership, follow through, and are proactive in addressing issues that need resolution.

• You have a passion for data and SQL, and you are inquisitive about how real-time analytical systems function.

• Familiarity with Traefik, Varnish, Redis, Terraform, or Ansible is beneficial but not mandatory.

• You communicate effectively in writing.

• You utilize AI tools such as Claude Code, Cursor, ChatGPT, and others.

• You are fluent in both English and Spanish.

• You are willing to participate in on-call rotations.

• You are located within an EU timezone.


🏝️ Benefits

• 22 days of holiday per year (in addition to your birthday and public holidays).

• Flexibility to work from any location that suits you best.

• We offer up to €2,800 to assist you in establishing your home workspace.

People also viewed

Ole & Lena Digital7 hours ago

Observability Engineer – Site Reliability Engineer

US flagCalifornia, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)$90 – $100/hour
ApplyView job
SOFTETA7 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sólides7 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sólides7 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zact7 hours ago

Cloud DevSecOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
3Pillar Global7 hours ago

Senior DevOps Engineer – AWS, Python

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers