
Site Reliability Engineer – Core Streaming
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Canada.
• Take ownership of the reliability, scalability, and operational health of Kafka clusters in multi-cloud and hybrid settings.
• Develop and sustain automation for cluster operations, upgrades, capacity scaling, and incident recovery.
• Collaborate with engineering teams to facilitate new streaming use cases, provide guidance on best practices, and ensure the reliability of data pipelines.
• Diagnose intricate issues impacting data flow, performance, or stability.
• Lead investigations into root causes.
• Implement Kafka version upgrades and platform migrations with minimal disruption to essential services.
• Engage in on-call rotations utilizing a geographically distributed follow-the-sun model.
• Promote automation and self-service capabilities for deploying, upgrading, and scaling streaming infrastructure.
• Strong foundation in SRE or infrastructure engineering.
• Proficiency in infrastructure-as-code, particularly with Terraform.
• Familiarity with configuration management tools such as Puppet, Ansible, or similar.
• Experience with cloud platforms; AWS is preferred.
• Background in Linux operations.
• Hands-on experience with Kafka or similar technologies in a production environment at scale.
• Knowledge in cluster upgrades, migrations, and capacity planning.
• Programming skills in Python, Java, or a similar language.
• Excellent debugging and systems-thinking abilities across distributed systems.
• Experience with Apache Flink or other stream processing frameworks is a plus.
• Understanding of Kafka Client APIs, including Producer, Consumer, and Streams is advantageous.
• Experience in developing internal self-service tools or developer platforms is a bonus.
• Background in incident response and management is desirable.
• Fully remote work available throughout Canada.
• Follow-the-sun on-call model; ensuring no one is on-call 24/7.
• Support from managers, mentors, and teams.
• Comprehensive benefits package (details available in the posting).
• Reasonable accommodations provided for individuals with disabilities during the job application process.
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.