
Senior Site Reliability Engineer
Posted Aug 19

Posted Aug 19
This is a fully remote position, open to applicants in Canada.
• Design, develop, and maintain highly scalable infrastructure utilizing Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.
• Take ownership of the AWS cloud environment, focusing on cost efficiency, security best practices, and high availability.
• Manage and enhance Kubernetes cluster environments using Helm, ArgoCD, Istio, and Kustomize.
• Oversee and scale data streaming infrastructure with Confluent Cloud and Apache Kafka.
• Deploy, configure, and manage Redis for caching and real-time data processing.
• Implement and sustain monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty.
• Enable controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.
• Collaborate with software development teams to integrate new services, applications, and features into the existing infrastructure.
• Diagnose and resolve complex system-level challenges while ensuring optimal performance and maximizing uptime.
• Propel continuous improvement of automation tools, operational processes, and engineering methodologies.
• Stay informed about emerging SRE trends and tools, assisting in the adoption of relevant industry advancements and best practices.
• 5+ years of experience in a Senior Site Reliability Engineer position or comparable role, with a strong focus on cloud infrastructure management and automation.
• Proficiency in Infrastructure as Code with Terraform and Terragrunt for enterprise-scale deployments.
• In-depth knowledge of AWS, including the design, implementation, and maintenance of secure, scalable, and resilient cloud architectures.
• Significant hands-on experience with distributed data streaming using Confluent Cloud and Apache Kafka.
• Demonstrated experience with Redis for caching and Amazon RDS for relational database management.
• Familiarity with enterprise search and analytics platforms such as OpenSearch, Elasticsearch, and ChaosSearch.
• Expertise in designing and implementing monitoring and alerting infrastructure using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty.
• Practical experience with feature flag systems, including LaunchDarkly/PostHog for managed release management.
• Extensive experience in administering production-grade Kubernetes with Helm, ArgoCD, and Istio; working knowledge of Kustomize.
• Strong analytical skills, capable of troubleshooting complex systems in production environments.
• Excellent communication and collaboration skills, with experience in Agile settings.
• Are you authorized to work in Canada without any restrictions?
• Will you require sponsorship now or in the future to work in Canada?
• Equity participation available to employees globally, with program details varying by location and employment structure.
• Competitive Health, Vision, Dental, and Life Insurance plans for eligible employees in the US.
• Robust 401k plan for eligible employees in the US.
• Discretionary Time Off for eligible employees in the US.
• Other minor perks.
• International employees receive competitive benefits in line with local market standards and applicable country requirements.
Keyfactor
XBOX
Trumid
GSB Solutions
Get handpicked remote jobs straight to your inbox weekly.