
Senior Site Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Manage the Kubernetes environment across development, continuous integration, pre-production, and one customer-facing production region.
• Ensure the production region operates in accordance with established service level objectives (SLOs), which include capacity planning, upgrade strategies, patch management, backup and restore processes, disaster recovery drills, and participation in on-call rotations.
• Oversee incident response activities, encompassing detection, mitigation, customer impact evaluation, root cause analysis, and conducting blameless postmortems.
• Develop and maintain Helm charts and umbrella releases, focusing on version control, values management, and upgrade pathways.
• Take ownership of CI/CD pipelines from start to finish, including building, testing, image publishing, chart packaging, release management, and hotfix/backport procedures.
• Automate the environment bootstrapping and seeding process to enable engineers to deploy the full stack with a single command.
• Manage and troubleshoot components such as PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability tools.
• Create observability and diagnostic solutions through metrics, dashboards, alerts, logging, and audit access functionalities.
• Advise product teams and regional operators on deployment architecture, GPU and resource allocation, role-based access control (RBAC), networking, and failure scenarios.
• Validate upgrade and migration protocols while providing comprehensive runbooks for handover.
• Implement security measures and ensure tenant isolation through least-privilege access, secret management, certificate and TLS lifecycle management, scanning, and maintaining audit trails.
• Promote infrastructure as code practices and ensure repeatability in processes.
• Mentor engineers on Kubernetes and operational methodologies through documentation and review sessions.
• A minimum of 10 years’ experience in DevOps, Site Reliability Engineering (SRE), platform, or infrastructure engineering, including responsibility for customer-facing Kubernetes environments.
• Expert knowledge of Kubernetes, encompassing workloads, networking, storage, RBAC, resource management, custom resource definitions (CRDs) and operators, as well as cluster upgrades.
• Demonstrated ability to respond to incidents effectively under SLA pressure, including managing on-call schedules, escalation protocols, postmortems, and corrective actions.
• Proficient in CI/CD engineering practices, including pipelines as code, reproducible builds, and artifact and release management.
• Experience with GitHub Actions or similar tools.
• Strong skills in scripting and automation.
• Familiarity with Go enough to read service code, trace issues, and report accurate bugs.
• Experience managing relational databases, identity providers, gateways, and message brokers within Kubernetes, including backup, restore, and upgrade processes.
• Technical consulting experience with well-documented runbooks, design insights, and incident reports across different global time zones.
• Strong proficiency in written English.
• Experience with Kubernetes at scale, Cluster API, controllers/operators, Docker, and the management of Helm charts throughout their lifecycle.
• Familiarity with container registries, versioned release, and backport workflows.
• Skills in Terraform, Ansible, or similar tools, along with GitOps solutions such as Argo CD or Flux.
• Knowledge of Keycloak and API gateway operations, including routing, plugins, TLS, and rate limiting.
• Experience with PostgreSQL operations and migrations, as well as Kafka or similar message-broker platforms.
• Proficiency in using Prometheus, Grafana, along with centralized logging and alerting that aligns with SLOs.
• Familiarity with AWS networking, IAM, load balancing, and managed Kubernetes services.
• A bachelor’s degree in Computer Science & Engineering or a related field, or 10 years of relevant experience.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company-sponsored outings, happy hours, hackathons, and tech discussions.
• Competitive salary along with a robust benefits package.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.