
Site Reliability Engineering Team Lead – Principal SRE
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in Canada, +1 more country.
• Lead Cerence's Site Reliability Engineering team, ensuring the reliability, availability, and operational health of its cloud-native automotive AI platform.
• Assist in selecting, mentoring, and technically developing the team across various locations.
• Establish technical direction and priorities, while providing performance and growth input to team managers.
• Design and maintain a sustainable on-call rotation, monitoring page load and team health.
• Own and advance the reliability roadmap over a 2–3 quarter timeframe.
• Define and govern SLI/SLO/SLA frameworks to achieve customer program availability targets of up to 99.95%.
• Act as the Tier 2 technical escalation point for major incidents in collaboration with the Global Operations Center.
• Promote a blameless postmortem culture and ensure actionable outcomes are derived.
• Lead and enhance Production Readiness / NFR reviews with development teams.
• Contribute to root cause analysis and take ownership of systemic improvements.
• Approve high-risk and out-of-window production changes.
• Set strategic direction for metrics, dashboards, alerting, escalation, and automation processes.
• Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks.
• Collaborate with DevOps and platform teams to enhance shared infrastructure.
• Integrate reliability into the SDLC by partnering with development managers and architects.
• Engage in service reliability consulting and architectural reviews.
• Communicate the reliability posture and associated risks to both technical and non-technical stakeholders.
• 8+ years of practical experience in site reliability, DevOps, or cloud platform roles, including experience in leading a team or overseeing a function.
• Proven ability to set technical direction and maintain standards across a team, with or without formal authority.
• Hands-on experience with container orchestration frameworks such as Kubernetes, Docker, and Istio.
• Familiarity with public cloud platforms, with a primary focus on Azure; experience with AWS and Google Cloud is also beneficial.
• Knowledge of observability tools, including metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana).
• Experience with CI/CD pipelines and infrastructure-as-code methodologies (e.g., Terraform, Flux).
• Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.
• Strong background in UNIX/Linux, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS).
• Exceptional written and verbal communication skills in English.
• Previous experience in site reliability leadership roles.
• Experience in leading distributed or multi-site technical teams.
• Background in high-availability service design, including redundancy, failover, and blast radius considerations.
• Familiarity with log aggregation and analytics platforms (e.g., Loki, Thanos).
• Understanding of ITSM and project management tools (e.g., Jira, Confluence).
• Experience working in automotive, embedded, or latency-sensitive production environments.
• Annual bonus opportunity.
• Comprehensive insurance coverage, including medical, dental, vision, life, and disability.
• Paid time off.
• Paid holidays.
• Company contributions to the RRSP (Registered Retirement Savings Plan).
• Equity awards available for certain roles and levels.
• Options for remote and/or hybrid work based on the position.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.