Site Reliability Engineer – Core Streaming

atYelpRemoteCA flagCanadaFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSeniorC$135k – C$185k/year

Posted Aug 19

This is a fully remote position, open to applicants in Canada.

📋 Description

• Take ownership of the reliability, scalability, and operational health of Kafka clusters in multi-cloud and hybrid settings.

• Develop and sustain automation for cluster operations, upgrades, capacity scaling, and incident recovery.

• Collaborate with engineering teams to facilitate new streaming use cases, provide guidance on best practices, and ensure the reliability of data pipelines.

• Diagnose intricate issues impacting data flow, performance, or stability.

• Lead investigations into root causes.

• Implement Kafka version upgrades and platform migrations with minimal disruption to essential services.

• Engage in on-call rotations utilizing a geographically distributed follow-the-sun model.

• Promote automation and self-service capabilities for deploying, upgrading, and scaling streaming infrastructure.


⛳️ Requirements

• Strong foundation in SRE or infrastructure engineering.

• Proficiency in infrastructure-as-code, particularly with Terraform.

• Familiarity with configuration management tools such as Puppet, Ansible, or similar.

• Experience with cloud platforms; AWS is preferred.

• Background in Linux operations.

• Hands-on experience with Kafka or similar technologies in a production environment at scale.

• Knowledge in cluster upgrades, migrations, and capacity planning.

• Programming skills in Python, Java, or a similar language.

• Excellent debugging and systems-thinking abilities across distributed systems.

• Experience with Apache Flink or other stream processing frameworks is a plus.

• Understanding of Kafka Client APIs, including Producer, Consumer, and Streams is advantageous.

• Experience in developing internal self-service tools or developer platforms is a bonus.

• Background in incident response and management is desirable.


🏝️ Benefits

• Fully remote work available throughout Canada.

• Follow-the-sun on-call model; ensuring no one is on-call 24/7.

• Support from managers, mentors, and teams.

• Comprehensive benefits package (details available in the posting).

• Reasonable accommodations provided for individuals with disabilities during the job application process.

People also viewed

Karat13 hours ago

Senior DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM1 day ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Jobsity1 day ago

Senior Site Reliability Engineer, AWS Multi-region

AR flagArgentina, +2 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM1 day ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Charger Logistics Inc.1 day ago

Site Reliability Engineer

IN flagIndia OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dev.Pro1 day ago

Senior DevOps Engineer

BR flagBrazil, +2 more countriesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers