
Staff Site Reliability Engineer
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Canada.
• Lead initiatives in reliability engineering and operational excellence for mission-critical services deployed on AWS and Kubernetes.
• Design, implement, and refine deployment, release, and rollback strategies across distributed systems.
• Establish CI/CD pipelines that prioritize security by default, incorporating automation, governance, and policy-driven controls.
• Enhance platform observability through metrics, logging, tracing, and actionable alerting systems.
• Define and advance Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards.
• Direct high-severity incident response efforts and conduct post-incident reviews.
• Collaborate with engineering teams to elevate platform standards, service resilience, runtime performance, and reliability practices.
• Mentor engineers on cloud-native technologies, Site Reliability Engineering (SRE) principles, and operational excellence.
• Design, build, and evolve foundational systems, tools, and operational practices.
• Architect scalable distributed systems while optimizing Kubernetes and AWS infrastructure.
• Minimize operational toil, enhance system performance, boost platform reliability, and facilitate business growth.
• Over 8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or similar cloud-native engineering roles.
• Extensive knowledge of AWS, including services such as EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security.
• Advanced experience in operating and scaling production Kubernetes environments.
• Strong practical experience with Istio service mesh.
• Demonstrated expertise in Infrastructure as Code, preferably utilizing AWS CDK.
• Experience in building and managing CI/CD pipelines with GitHub Actions or equivalent platforms.
• Robust troubleshooting, performance optimization, and incident management skills in distributed systems.
• Excellent communication, collaboration, and technical leadership capabilities.
• Experience in designing and operating monitoring, logging, tracing, and alerting solutions.
• Familiarity with AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tools.
• Experience in defining SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.
• Proficiency in TypeScript and Node.js.
• Experience in developing scalable backend services, APIs, and event-driven systems.
• Understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.
• Experience in implementing zero-trust architectures, mTLS, and service-to-service security measures.
• Familiarity with automated testing, code reviews, and observability-driven development practices.
• Understanding of resilience engineering, autoscaling, disruption management, failure testing, and safe deployment strategies.
• Experience in progressive delivery, working in regulated or security-sensitive SaaS environments, knowledge of FinOps, internal developer platforms, or cloud-native certifications is a plus.
• Flexible work options.
• Remote work opportunities.
• Generous time-off policies.
• Competitive salary.
• Health insurance.
• Retirement plans.
• Recognition programs.
• Performance bonuses.
• Opportunities for career growth.
• International projects.
• Collaboration with a diverse, global team.
• An inclusive and equitable workplace.
• Accommodations available during the application or interview process.
Ninja - نينجا
Dreamix
Scalingo
inDrive
Get handpicked remote jobs straight to your inbox weekly.