Remotery

Senior Site Reliability Engineer

Posted Jul 14

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, develop, and enhance Kubernetes infrastructure to support secure, multi-tenant, high-availability applications.

• Establish and manage AI tooling infrastructure by deploying MCP servers and ensuring secure, governed access to AI for production systems.

• Optimize and sustain CI/CD pipelines, enhancing their reliability, speed, and rollback capabilities.

• Implement advanced delivery strategies, including blue/green and canary deployments.

• Promote Infrastructure as Code practices utilizing Terraform, Helm, and Argo CD, creating reusable patterns for the organization.

• Manage and refine streaming and analytics infrastructure, specifically Kafka, Flink, and ClickHouse.

• Integrate automated testing within the CI/CD lifecycle.

• Enhance system observability by defining SLOs, alerts, and dashboards.

• Lead incident response and postmortem analyses, concentrating on root cause identification and sustainable solutions.

• Mentor engineers across various teams on Kubernetes, CI/CD, and cloud infrastructure.


⛳️ Requirements

• Over 6 years of experience in SRE, DevOps, or Infrastructure positions, with substantial hands-on experience in production Kubernetes environments.

• Practical experience incorporating AI/LLM tools into engineering or operational processes (e.g., MCP servers, AI agents managing infrastructure), along with a solid understanding of the security and governance implications of providing AI access to production.

• Demonstrated success in constructing CI/CD pipelines (using GitHub Actions, Jenkins, GitLab CI, or similar tools).

• Strong knowledge of Kubernetes internals and managed services such as EKS, GKE, or AKS.

• Proficiency in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps methodologies.

• Skilled in programming languages such as Python, Bash, or Go.

• Familiarity with observability tools like Prometheus, Grafana, Datadog, and OpenTelemetry.

• Production-level experience with Kafka, Flink, and ClickHouse.

• Excellent communication and collaboration skills across teams.


🏝️ Benefits

• Competitive salary

• Stock options

• Health benefits

• Unlimited PTO

• Parental leave

• Tuition reimbursements

People also viewed

BeyondTrust1 day ago

Senior DevOps Engineer

CA flagCanada OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ontrac Solutions1 day ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath1 day ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions1 day ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA1 day ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro1 day ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers