Senior Site Reliability Engineer

Posted Aug 28

This is a fully remote position, open to applicants in Canada.

📋 Description

• Develop and manage the infrastructure that supports a large-scale sports betting and media platform.

• Take ownership of essential infrastructure encompassing compute, networking, storage, and cloud services.

• Lead intricate infrastructure migrations and projects across production settings and various jurisdictions.

• Create and sustain platform tooling and automation utilizing ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows.

• Provide infrastructure consulting, resolve dependencies, conduct architecture reviews, and facilitate platform-tool adoption for development teams.

• Design and enhance Datadog observability, alerting, dashboards, and runbooks.

• Deliver operational support and incident response through systematic debugging and root cause analysis.

• Engage in on-call rotations as required.

• Mentor colleagues and participate in architectural decision-making and ongoing improvement initiatives.


⛳️ Requirements

• Over 5 years of experience in a comparable position (DevOps, Site Reliability Engineer).

• Extensive experience in operating and troubleshooting Kubernetes within a production Linux environment (including cluster lifecycle, networking, storage, and scheduling).

• Familiarity with AWS, GCP, and/or on-premise environments.

• Proficiency in at least two programming languages: Go, Python, Bash/Shell.

• Comprehensive understanding of distributed systems, failure modes, networking basics, capacity planning, and performance evaluation.

• Experience with GitOps and CI/CD workflows (such as ArgoCD, Helm, GitHub Actions, or equivalent).

• Knowledge of infrastructure-as-code tools (like Terraform, Helm, or similar).

• Proven track record of leading complex migrations or infrastructure projects involving cross-team dependencies.

• Strong skills in incident response and troubleshooting.

• Effective technical communication and documentation abilities.

• Preferred experience with service mesh technologies (such as Istio, Cilium).

• Familiarity with distributed storage systems (like Ceph or similar) is preferred.

• Experience with bare-metal Kubernetes or Talos OS is preferred.

• Exposure to regulated environments is preferred.

• Experience with Datadog or equivalent observability platforms at scale is preferred.

• Familiarity with PostgreSQL, PgBouncer, or database migration tools is preferred.


🏝️ Benefits

• Competitive compensation package.

• Comprehensive benefits package.

• Enjoy a fun and relaxed work environment.

• Receive reimbursements for education and conferences.

• Eligibility for bonuses for most non-sales roles.

• Access to best-in-class benefits with tailored support options for physical, financial, and emotional well-being.

People also viewed

knowmad mood1 day ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions2 days ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services2 days ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services2 days ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal2 days ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios2 days ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers