
Site Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Canada.
• Evaluate the maturity of services and offer valuable insights to development teams.
• Collaborate with development teams to instill observability best practices.
• Empower development teams to independently manage their service deployment, support, and infrastructure.
• Guide developers on reliability practices, aiming to foster self-sufficiency.
• Serve as the liaison, advocate, and observer for the Platform Division teams to enhance the adoption of tools and practices among development teams.
• Extensive knowledge of observability practices in distributed system environments and their impact on system design and team dynamics.
• Hands-on experience with SRE principles (SLOs, error budgets, incident management).
• 3–5+ years of experience in software development, SRE, DevOps, or production development roles, particularly in operating production systems.
• Skilled in cloud-native platforms and infrastructure-as-code tools and concepts.
• Familiarity with at least one programming language (experience with TypeScript/Node.js is advantageous).
• Strong communication and collaboration skills across both technical and non-technical teams.
• Ability to simplify complex reliability concepts into actionable insights.
• Passion for empowering teams to thrive independently, gauging success by the decrease in dependency on your support.
• Competitive salary and substantial equity opportunities.
• Comprehensive healthcare, dental, and vision plans.
• Enrollment in a 401(k) / RRSP program.
• Flexible PTO policy allowing you to take what you need.
• A work culture that promotes:
• Collaboration with individuals worldwide who embody the MaintainX values: Smart Humble Optimist.
• A belief in meritocracy, where contributions and ideas are recognized and celebrated.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.