Senior Site Reliability Engineer

Posted Aug 19

This is a fully remote position, open to applicants in Canada.

📋 Description

• Design, develop, and maintain highly scalable infrastructure utilizing Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.

• Take ownership of the AWS cloud environment, focusing on cost efficiency, security best practices, and high availability.

• Manage and enhance Kubernetes cluster environments using Helm, ArgoCD, Istio, and Kustomize.

• Oversee and scale data streaming infrastructure with Confluent Cloud and Apache Kafka.

• Deploy, configure, and manage Redis for caching and real-time data processing.

• Implement and sustain monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty.

• Enable controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.

• Collaborate with software development teams to integrate new services, applications, and features into the existing infrastructure.

• Diagnose and resolve complex system-level challenges while ensuring optimal performance and maximizing uptime.

• Propel continuous improvement of automation tools, operational processes, and engineering methodologies.

• Stay informed about emerging SRE trends and tools, assisting in the adoption of relevant industry advancements and best practices.


⛳️ Requirements

• 5+ years of experience in a Senior Site Reliability Engineer position or comparable role, with a strong focus on cloud infrastructure management and automation.

• Proficiency in Infrastructure as Code with Terraform and Terragrunt for enterprise-scale deployments.

• In-depth knowledge of AWS, including the design, implementation, and maintenance of secure, scalable, and resilient cloud architectures.

• Significant hands-on experience with distributed data streaming using Confluent Cloud and Apache Kafka.

• Demonstrated experience with Redis for caching and Amazon RDS for relational database management.

• Familiarity with enterprise search and analytics platforms such as OpenSearch, Elasticsearch, and ChaosSearch.

• Expertise in designing and implementing monitoring and alerting infrastructure using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty.

• Practical experience with feature flag systems, including LaunchDarkly/PostHog for managed release management.

• Extensive experience in administering production-grade Kubernetes with Helm, ArgoCD, and Istio; working knowledge of Kustomize.

• Strong analytical skills, capable of troubleshooting complex systems in production environments.

• Excellent communication and collaboration skills, with experience in Agile settings.

• Are you authorized to work in Canada without any restrictions?

• Will you require sponsorship now or in the future to work in Canada?


🏝️ Benefits

• Equity participation available to employees globally, with program details varying by location and employment structure.

• Competitive Health, Vision, Dental, and Life Insurance plans for eligible employees in the US.

• Robust 401k plan for eligible employees in the US.

• Discretionary Time Off for eligible employees in the US.

• Other minor perks.

• International employees receive competitive benefits in line with local market standards and applicable country requirements.

People also viewed

Keyfactor1 day ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
XBOX1 day ago

Senior Site Reliability Engineer, Data & Analytics

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$102.8k – $190.2k/year
ApplyView job
Trumid1 day ago

Senior Database Reliability Engineer – DBRE

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$225k – $265k/year
ApplyView job
GSB Solutions1 day ago

Senior Release Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$80k/year
ApplyView job
Smartcat1 day ago

Head of Infrastructure and DevOps

PT flagPortugal OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Smartcat1 day ago

Head of Infrastructure and DevOps

GE flagGeorgia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers