Remotery

Site Reliability Engineer, Core Streaming

Posted Jul 17

This is a fully remote position, open to applicants in California, +3 more states.

📋 Description

• Take responsibility for the reliability, scalability, and operational health of Kafka clusters in multi-cloud and hybrid environments.

• Develop and sustain automation for cluster operations, upgrades, capacity scaling, and incident recovery.

• Collaborate with engineering teams to facilitate new streaming use cases, provide best practice guidance, and ensure the reliability of data pipelines.

• Diagnose complex issues impacting data flow, performance, or stability, and spearhead root cause analyses.

• Carry out Kafka version upgrades and platform migrations with minimal impact on critical services.

• Participate in on-call rotations. Our geographically distributed SRE teams utilize a “follow-the-sun” model, ensuring that no one is required to be on-call 24/7!


⛳️ Requirements

• Strong foundation in SRE or infrastructure engineering: experience with infrastructure-as-code (Terraform), configuration management (Puppet, Ansible, or a similar tool), cloud platforms (AWS preferred), and Linux operations.

• Hands-on experience with Kafka or similar technologies at scale, including cluster upgrades, migrations, and capacity planning.

• Proficiency in programming with Python, Java, or similar languages for tooling and automation purposes.

• Excellent debugging and systems-thinking abilities, capable of tracing data flow issues end-to-end across distributed systems.

• Nice to Have

• Experience with Apache Flink or other stream processing frameworks.

• Familiarity with Kafka Client APIs (Producer, Consumer, Streams).

• Experience in developing internal self-service tools or developer platforms.

• Background in incident response and management.


🏝️ Benefits

• There may be flexibility with the range included in this posting should a candidate be leveled higher or lower than the posted range. This opportunity offers the option to work fully remotely from any location in the US. More information about Yelp's five-star benefits can be found here!

People also viewed

The Codest4 days ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUM4 days ago

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
Sólides4 days ago

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Resilinc4 days ago

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group4 days ago

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLER4 days ago

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers