
Engineering Manager – Site Reliability
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Canada.
• Oversee, mentor, and cultivate a team of Site Reliability Engineers.
• Establish technical direction and priorities focusing on reliability, scalability, availability, performance, and operational readiness.
• Collaborate with engineering, product, security, infrastructure, and other cross-functional teams.
• Define reliability standards, influence system design, and implement initiatives that enhance customer and developer experiences.
• Lead incident management, incident response, post-incident analysis, service-level objectives, capacity planning, observability, and ongoing risk mitigation.
• Advocate for automation and self-service tools to minimize operational toil and boost deployment confidence.
• Balance immediate operational requirements with long-term investments while making trade-offs in a rapidly growing environment.
• Communicate reliability risks, project updates, trade-offs, and decisions to both technical and non-technical stakeholders.
• Make informed decisions with incomplete information, respond effectively during incidents, and guide teams in learning from failures.
• Bachelor’s degree in Computer Science, Engineering, or a related technical discipline, or equivalent hands-on experience.
• A minimum of seven years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field.
• At least two years of experience managing, mentoring, or leading engineering teams.
• Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or similar production technologies.
• Background in leading or participating in production incident response, post-incident evaluations, reliability enhancement initiatives, and operational readiness practices.
• Experience in articulating technical risks, priorities, and trade-offs to engineering leaders and cross-functional stakeholders.
• Preferred: experience in leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams.
• Preferred: experience in operating highly available services at scale and enhancing service-level objectives, observability, capacity, or disaster recovery capabilities.
• Preferred: familiarity with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation.
• Preferred: experience in building or evolving reliability programs across multiple engineering teams.
• Preferred: proven ability to foster alignment across teams, navigate uncertainty, and transform complex technical challenges into clear, actionable plans.
• Preferred: leadership style rooted in empathy, transparency, direct communication, collaboration, and inclusive team development.
• Flexibility to work from home, an office, or a coffee shop.
• Regular in-person events.
• New hire equity grant.
• Annual refresh grants.
• Competitive compensation and benefits.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.