
Senior Site Reliability Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Canada.
• Develop and manage the infrastructure that supports a large-scale sports betting and media platform.
• Take ownership of essential infrastructure encompassing compute, networking, storage, and cloud services.
• Lead intricate infrastructure migrations and projects across production settings and various jurisdictions.
• Create and sustain platform tooling and automation utilizing ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows.
• Provide infrastructure consulting, resolve dependencies, conduct architecture reviews, and facilitate platform-tool adoption for development teams.
• Design and enhance Datadog observability, alerting, dashboards, and runbooks.
• Deliver operational support and incident response through systematic debugging and root cause analysis.
• Engage in on-call rotations as required.
• Mentor colleagues and participate in architectural decision-making and ongoing improvement initiatives.
• Over 5 years of experience in a comparable position (DevOps, Site Reliability Engineer).
• Extensive experience in operating and troubleshooting Kubernetes within a production Linux environment (including cluster lifecycle, networking, storage, and scheduling).
• Familiarity with AWS, GCP, and/or on-premise environments.
• Proficiency in at least two programming languages: Go, Python, Bash/Shell.
• Comprehensive understanding of distributed systems, failure modes, networking basics, capacity planning, and performance evaluation.
• Experience with GitOps and CI/CD workflows (such as ArgoCD, Helm, GitHub Actions, or equivalent).
• Knowledge of infrastructure-as-code tools (like Terraform, Helm, or similar).
• Proven track record of leading complex migrations or infrastructure projects involving cross-team dependencies.
• Strong skills in incident response and troubleshooting.
• Effective technical communication and documentation abilities.
• Preferred experience with service mesh technologies (such as Istio, Cilium).
• Familiarity with distributed storage systems (like Ceph or similar) is preferred.
• Experience with bare-metal Kubernetes or Talos OS is preferred.
• Exposure to regulated environments is preferred.
• Experience with Datadog or equivalent observability platforms at scale is preferred.
• Familiarity with PostgreSQL, PgBouncer, or database migration tools is preferred.
• Competitive compensation package.
• Comprehensive benefits package.
• Enjoy a fun and relaxed work environment.
• Receive reimbursements for education and conferences.
• Eligibility for bonuses for most non-sales roles.
• Access to best-in-class benefits with tailored support options for physical, financial, and emotional well-being.
knowmad mood
RealTime eClinical Solutions
Koniag Government Services
Koniag Government Services
Get handpicked remote jobs straight to your inbox weekly.