Remotery

Senior Site Reliability Engineer

Posted 4 days ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Design, develop, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.

• Construct and manage AI tooling infrastructure — establish MCP servers and create secure, governed AI access and guardrails for production systems.

• Enhance and maintain CI/CD pipelines, boosting reliability, speed, and rollback safety.

• Implement advanced delivery strategies such as blue/green and canary deployments.

• Promote Infrastructure as Code with Terraform, Helm, and Argo CD, defining reusable patterns for the organization.

• Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and ClickHouse.

• Integrate automated testing into the CI/CD lifecycle.

• Enhance system observability — establish SLOs, alerts, and dashboards.

• Lead incident response and postmortems, concentrating on root causes and long-term solutions.

• Mentor engineers across teams on Kubernetes, CI/CD, and cloud infrastructure.


⛳️ Requirements

• Over 6 years of experience in SRE, DevOps, or Infrastructure roles, with substantial production Kubernetes expertise.

• Practical experience integrating AI/LLM tooling into engineering or operational workflows (e.g., MCP servers, AI agents interacting with infrastructure), along with a solid understanding of the security and governance implications of providing AI access to production.

• Demonstrated success in building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, or similar).

• Strong knowledge of Kubernetes internals and managed services like EKS, GKE, or AKS.

• Expertise in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps methodologies.

• Proficient in Python, Bash, or Go programming languages.

• Familiarity with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry).

• Hands-on experience with Kafka, Flink, and ClickHouse in production environments.

• Excellent communication and cross-team collaboration skills.


🏝️ Benefits

• Competitive salary

• Stock options

• Health benefits

• Unlimited PTO

• Parental leave

• Tuition reimbursements

People also viewed

Fundraise Up16 hours ago

Senior DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€6,000 – €6,800/month
ApplyView job
Empower16 hours ago

Lead Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$114k – $165.3k/year
ApplyView job
Harrods16 hours ago

DevOps Manager

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aufinity Group | España16 hours ago

Software Developer – DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zipdev16 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Valtech17 hours ago

Senior Site Reliability Engineer

MK flagMacedonia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers