Site Reliability Engineering Team Lead – Principal SRE

Posted Sep 17

This is a fully remote position, open to applicants in Canada, +1 more country.

📋 Description

• Lead Cerence's Site Reliability Engineering team, ensuring the reliability, availability, and operational health of its cloud-native automotive AI platform.

• Assist in selecting, mentoring, and technically developing the team across various locations.

• Establish technical direction and priorities, while providing performance and growth input to team managers.

• Design and maintain a sustainable on-call rotation, monitoring page load and team health.

• Own and advance the reliability roadmap over a 2–3 quarter timeframe.

• Define and govern SLI/SLO/SLA frameworks to achieve customer program availability targets of up to 99.95%.

• Act as the Tier 2 technical escalation point for major incidents in collaboration with the Global Operations Center.

• Promote a blameless postmortem culture and ensure actionable outcomes are derived.

• Lead and enhance Production Readiness / NFR reviews with development teams.

• Contribute to root cause analysis and take ownership of systemic improvements.

• Approve high-risk and out-of-window production changes.

• Set strategic direction for metrics, dashboards, alerting, escalation, and automation processes.

• Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks.

• Collaborate with DevOps and platform teams to enhance shared infrastructure.

• Integrate reliability into the SDLC by partnering with development managers and architects.

• Engage in service reliability consulting and architectural reviews.

• Communicate the reliability posture and associated risks to both technical and non-technical stakeholders.


⛳️ Requirements

• 8+ years of practical experience in site reliability, DevOps, or cloud platform roles, including experience in leading a team or overseeing a function.

• Proven ability to set technical direction and maintain standards across a team, with or without formal authority.

• Hands-on experience with container orchestration frameworks such as Kubernetes, Docker, and Istio.

• Familiarity with public cloud platforms, with a primary focus on Azure; experience with AWS and Google Cloud is also beneficial.

• Knowledge of observability tools, including metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana).

• Experience with CI/CD pipelines and infrastructure-as-code methodologies (e.g., Terraform, Flux).

• Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.

• Strong background in UNIX/Linux, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS).

• Exceptional written and verbal communication skills in English.

• Previous experience in site reliability leadership roles.

• Experience in leading distributed or multi-site technical teams.

• Background in high-availability service design, including redundancy, failover, and blast radius considerations.

• Familiarity with log aggregation and analytics platforms (e.g., Loki, Thanos).

• Understanding of ITSM and project management tools (e.g., Jira, Confluence).

• Experience working in automotive, embedded, or latency-sensitive production environments.


🏝️ Benefits

• Annual bonus opportunity.

• Comprehensive insurance coverage, including medical, dental, vision, life, and disability.

• Paid time off.

• Paid holidays.

• Company contributions to the RRSP (Registered Retirement Savings Plan).

• Equity awards available for certain roles and levels.

• Options for remote and/or hybrid work based on the position.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers