Site Reliability Engineer

atRentsyncRemoteCA flagCanadaFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSeniorC$80k – C$110k/year

Posted 5 days ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Serve as the initial point of contact for production alerts and incidents across various services, managing the process from triage to resolution.

• Identify and resolve issues directly in AWS and Kubernetes, addressing problems such as failing pods, resource exhaustion, deployment errors, networking/DNS, database, and caching issues.

• Execute rollbacks, scale, reconfigure, or apply patches to infrastructure to promptly restore services.

• Escalate issues to development teams only when a code modification is required, ensuring a detailed diagnosis is provided.

• Oversee the PagerDuty configuration and incident response during business hours while striving to minimize MTTD and MTTR.

• Conduct blameless post-mortems and lead technical follow-up initiatives.

• Automate runbooks and repetitive operational tasks, utilizing AI-assisted triage, investigation, and remediation.

• Develop and maintain monitoring solutions for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry.

• Monitor production releases and detect regressions in latency, errors, or resource consumption.

• Design and manage synthetic checks, smoke tests, health checks, and load/performance tests.

• Collaborate with engineering teams to address performance and reliability challenges, defining SLOs, SLIs, and error budgets.

• Strengthen the platform through Terraform modifications, Kubernetes resource optimization, autoscaling, CI/CD checks, secrets management, and IAM enhancements.

• Maintain comprehensive service documentation and document architectural decisions.


⛳️ Requirements

• Minimum of 3 years in a cloud engineering, DevOps, or SRE position providing support for production web applications.

• Practical experience in responding to production incidents, including diagnosing and resolving issues directly.

• Strong hands-on experience with AWS in production environments, including EKS, EC2, RDS, VPC networking, IAM, and CloudWatch.

• Extensive experience with Kubernetes in production, encompassing troubleshooting, debugging, and workload monitoring.

• Proven track record of identifying and resolving performance and reliability issues in collaboration with engineering teams.

• Familiarity with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch.

• Experience with on-call and alerting tools like PagerDuty.

• Capability to develop automated tests or checks for production reliability, including synthetic, smoke, health, or load tests.

• Willingness to partake in a potential future on-call rotation outside of regular hours.

• Proficiency in infrastructure as code using Terraform and experience working within CI/CD pipelines.

• Strong grounding in Linux, networking, and container fundamentals.

• Competency in scripting and automation using Bash, Python, or similar languages.

• Ability to communicate calmly and clearly during incidents and across teams.

• Preferred: Familiarity with AI tools for SRE tasks; experience with Azure and possibly GCP; knowledge of multiple technology stacks; AWS certification; proficiency with LGTM or OpenTelemetry at scale; experience with k6, Locust, or JMeter; MySQL/PostgreSQL management; Redis/Memcached optimization; Cloudflare; and expertise in cloud cost management and capacity planning.


🏝️ Benefits

• Remote work opportunity.

• Preferential consideration may be offered to candidates located within a reasonable commuting distance to one of the offices.

• Equal opportunity employer.

• Unique accommodations available during the interview process.

• Potential criminal background check during the final interview phase.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers