Remotery

Senior Site Reliability Engineer, DevEx

Posted 4 hours ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Design and create the foundational infrastructure elements that dictate how our CI/CD platform, build systems, and developer environments scale across the entire engineering organization.

• Assist in the development and management of the Kubernetes-based control plane that supports our CI/CD platform, which includes:

• - Infrastructure for GitHub Actions self-hosted runners (including autoscaling, isolation, and cost/performance tuning)

• - GitHub Apps and GitHub-as-code (covering permissions, webhooks, and automation across the organization)

• - Secure network access for CI/CD and remote development environments utilizing Tailscale

• - GitOps-driven deployment of platform services through Flux

• - Creation of ephemeral and on-demand developer environments and build systems

• Develop the essential infrastructure components — such as Kubernetes Operators and automation for scaling — that product teams will adopt directly, thereby minimizing tailored CI/CD and environment tools for each team.

• Construct the systems that outline how engineering teams build, test, and deploy, influencing the reliability and scalability of the developer experience across the organization.


⛳️ Requirements

• 6–9+ years of experience in SRE / Platform / Infrastructure Engineering

• Demonstrated experience in scaling Kubernetes within high-throughput production environments

• In-depth knowledge of Kubernetes beyond cluster operations — including internals, scheduler behavior, custom resources, and diagnosing cluster-scale failures

• Experience in building platform infrastructure, control planes, or Kubernetes Operators (rather than merely utilizing them)

• Strong background in distributed systems and ensuring production reliability

• Ownership of Terraform/GitOps — designing and managing automation rather than simply executing playbooks

• Familiarity with GitOps workflows (Flux / ArgoCD)

• Practical experience with CI/CD platforms at scale: GitHub Actions (self-hosted runners, workflows-as-code), GitHub Apps, and build systems

• Production experience with AWS/cloud infrastructure

• Proficiency in Go (strongly preferred) or another systems programming language

• Proven history of constructing infrastructure primitives instead of mainly providing support/operations, demonstrating an automation-first approach.


🏝️ Benefits

• Comprehensive benefits

• Long-term incentives

People also viewed

Quantiphi4 hours ago

DevOps Engineer, Observability

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NIR-YU4 hours ago

Professional Cloud DevOps Engineer – Google Cloud, Certificado

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Bet On Talent4 hours ago

Senior DevSecOps Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zignaly4 hours ago

Infrastructure / Systems Operations Engineer

AE flagUnited Arab Emirates (UAE) OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$40k – $50k/year
ApplyView job
knowmad mood4 hours ago

Senior Business Consultant – DevOps, Atlassian

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Astronomer4 hours ago

Customer Reliability Engineer, Airflow

US flagCalifornia, +8 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$125k – $130k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers