
Senior Site Reliability Engineer
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in United States.
• Design, develop, and enhance Kubernetes infrastructure to support secure, multi-tenant, high-availability applications.
• Establish and manage AI tooling infrastructure by deploying MCP servers and ensuring secure, governed access to AI for production systems.
• Optimize and sustain CI/CD pipelines, enhancing their reliability, speed, and rollback capabilities.
• Implement advanced delivery strategies, including blue/green and canary deployments.
• Promote Infrastructure as Code practices utilizing Terraform, Helm, and Argo CD, creating reusable patterns for the organization.
• Manage and refine streaming and analytics infrastructure, specifically Kafka, Flink, and ClickHouse.
• Integrate automated testing within the CI/CD lifecycle.
• Enhance system observability by defining SLOs, alerts, and dashboards.
• Lead incident response and postmortem analyses, concentrating on root cause identification and sustainable solutions.
• Mentor engineers across various teams on Kubernetes, CI/CD, and cloud infrastructure.
• Over 6 years of experience in SRE, DevOps, or Infrastructure positions, with substantial hands-on experience in production Kubernetes environments.
• Practical experience incorporating AI/LLM tools into engineering or operational processes (e.g., MCP servers, AI agents managing infrastructure), along with a solid understanding of the security and governance implications of providing AI access to production.
• Demonstrated success in constructing CI/CD pipelines (using GitHub Actions, Jenkins, GitLab CI, or similar tools).
• Strong knowledge of Kubernetes internals and managed services such as EKS, GKE, or AKS.
• Proficiency in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps methodologies.
• Skilled in programming languages such as Python, Bash, or Go.
• Familiarity with observability tools like Prometheus, Grafana, Datadog, and OpenTelemetry.
• Production-level experience with Kafka, Flink, and ClickHouse.
• Excellent communication and collaboration skills across teams.
• Competitive salary
• Stock options
• Health benefits
• Unlimited PTO
• Parental leave
• Tuition reimbursements
BeyondTrust
Ontrac Solutions
CyberSheath
Ontrac Solutions
Get handpicked remote jobs straight to your inbox weekly.