
Senior Site Reliability Engineer
Posted Jul 20

Posted Jul 20
This is a fully remote position, open to applicants in Canada.
• Oversee the execution of infrastructure projects.
• Strategize and carry out maintenance tasks that involve higher risks.
• Assist in incident resolution and take part in an on-call schedule.
• Collaborate with product and software development teams to enhance the resilience and reliability of our offerings.
• Guide team members across all facets of Site Reliability Engineering (SRE) tasks.
• Effectively manage your productivity and workload while working remotely.
• Employ a data-driven methodology to identify necessary changes in the product architecture to boost reliability, performance, and availability.
• Gain a comprehensive understanding of production environments and the entire delivery lifecycle.
• Recognize components of the system that are not scalable and initiate solutions for these challenges.
• Uphold and enhance Service Level Indicators (SLI) that correspond with availability and performance objectives.
• Infuse quality into the team’s output by promoting refactoring, testing, and segmenting the team’s tasks into manageable, releasable units.
• Advocate for automation and ongoing improvements to diminish operational burdens and enhance platform reliability.
• Over 8 years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering roles supporting production cloud environments.
• More than 3 years in a senior or technical leadership role, with a proven track record of managing critical production systems and mentoring engineers.
• Practical experience with contemporary cloud platforms (AWS, GCP, or Azure), encompassing deployment, scaling, monitoring, and cost optimization of SaaS applications.
• At least 5 years of relevant experience in DevOps, SRE, or infrastructure engineering roles.
• Demonstrated success in implementing and managing observability stacks (e.g., Datadog, Prometheus, New Relic, Splunk) and enhancing SLIs/SLOs.
• Experience in incident management, including involvement in on-call rotations and leading post-incident reviews focused on continuous improvement.
• Familiarity with AI/LLM upskilling — adept at utilizing AI and agentic tools (e.g., Claude Code) to streamline investigation, automation, and delivery processes.
• Hands-on experience with service mesh technologies (Envoy/Istio) for traffic management, observability, and secure service-to-service communication.
• Experience with NATS or similar messaging/streaming systems for event-driven and distributed architectures.
• Purposeful Work: Contribute to global efforts in advancing sustainable food production.
• Our People: Collaborate with a fun, supportive, and dynamic team.
• Recharge: Enjoy a generous vacation policy, company-sponsored holidays, and a year-end winter break.
• Work Flexibility: Benefit from hybrid work arrangements and a strong culture of work-life balance.
• Prioritize Your Well-Being: Access extensive health plans tailored to support your physical and mental wellness.
• Group RRSP with a 3% company-paid match after three months of employment.
• Convenient office location accessible via public transit and bike paths.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.