
Senior Site Reliability Engineer, DevEx
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in Canada.
• Design and create the foundational infrastructure elements that dictate how our CI/CD platform, build systems, and developer environments scale across the entire engineering organization.
• Assist in the development and management of the Kubernetes-based control plane that supports our CI/CD platform, which includes:
• - Infrastructure for GitHub Actions self-hosted runners (including autoscaling, isolation, and cost/performance tuning)
• - GitHub Apps and GitHub-as-code (covering permissions, webhooks, and automation across the organization)
• - Secure network access for CI/CD and remote development environments utilizing Tailscale
• - GitOps-driven deployment of platform services through Flux
• - Creation of ephemeral and on-demand developer environments and build systems
• Develop the essential infrastructure components — such as Kubernetes Operators and automation for scaling — that product teams will adopt directly, thereby minimizing tailored CI/CD and environment tools for each team.
• Construct the systems that outline how engineering teams build, test, and deploy, influencing the reliability and scalability of the developer experience across the organization.
• 6–9+ years of experience in SRE / Platform / Infrastructure Engineering
• Demonstrated experience in scaling Kubernetes within high-throughput production environments
• In-depth knowledge of Kubernetes beyond cluster operations — including internals, scheduler behavior, custom resources, and diagnosing cluster-scale failures
• Experience in building platform infrastructure, control planes, or Kubernetes Operators (rather than merely utilizing them)
• Strong background in distributed systems and ensuring production reliability
• Ownership of Terraform/GitOps — designing and managing automation rather than simply executing playbooks
• Familiarity with GitOps workflows (Flux / ArgoCD)
• Practical experience with CI/CD platforms at scale: GitHub Actions (self-hosted runners, workflows-as-code), GitHub Apps, and build systems
• Production experience with AWS/cloud infrastructure
• Proficiency in Go (strongly preferred) or another systems programming language
• Proven history of constructing infrastructure primitives instead of mainly providing support/operations, demonstrating an automation-first approach.
• Comprehensive benefits
• Long-term incentives
Quantiphi
NIR-YU
Bet On Talent
Zignaly
Get handpicked remote jobs straight to your inbox weekly.