Remotery

Senior Site Reliability Engineer – Kubernetes

Posted 20 hours ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Assist in the deployment, operation, and maintenance of production services utilizing Kubernetes.

• Oversee service health and investigate incidents occurring in production across distributed applications.

• Engage in on-call support, incident response, root cause analysis, postmortem reviews, and enhancements in reliability.

• Diagnose runtime issues, networking challenges, and service-to-service problems in collaboration with engineering teams.

• Facilitate CI/CD, GitOps-based deployments, observability, and monitoring of production environments.

• Operate within a client-directed backlog and adhere to set priorities.


⛳️ Requirements

• A minimum of 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a similar role.

• Recent and substantial hands-on experience in supporting production services based on Kubernetes.

• Preferred: At least 3 years of hands-on experience with Kubernetes in production.

• Familiarity with Kubernetes production operations, including deployment, scaling, rollout/rollback, resource optimization, and troubleshooting service-to-service issues.

• Extensive experience in production incident response, covering on-call duties, runbooks, postmortems, and paging protocols.

• Proficiency in Splunk for log aggregation, search capabilities, and troubleshooting in production environments.

• Experience with Prometheus and Grafana, particularly in creating alerting rules and dashboards.

• Knowledge of CI/CD and infrastructure-as-code practices for containerized deployments, including tools like Helm and GitOps platforms such as ArgoCD or Flux.

• Solid understanding of Linux and networking fundamentals, covering DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking.

• Expertise in Node.js production troubleshooting, including heap snapshots, CPU profiling, event-loop blocking, memory management, worker/process isolation, and V8 isolates or similar runtime models.

• Proficiency in JVM/Java production troubleshooting, including GC log analysis, thread dump examination, JVM tuning, and investigating service latency in Java applications.

• Familiarity with in-memory caching systems like Redis/Valkey, including key design, TTL/eviction tuning, and cache invalidation strategies.

• Experience with service-to-service authentication mechanisms, such as mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication.

• Experience with Web Components/Lit is a plus.

• Knowledge of server-side rendering or isomorphic runtime environments is an advantage.

• Familiarity with canary rollout and multi-version production operations is desirable.

• Experience with distributed tracing and request-context correlation is beneficial.

• Knowledge of KEDA or event-driven autoscaling is a plus.

• Familiarity with enterprise platform integration layers is advantageous.


🏝️ Benefits

• Competitive salary and laptop provided.

• Opportunities for professional development and training.

• Engage with state-of-the-art cloud and container technologies.

• Flexible work arrangements and a collaborative team environment.

• Contribute to organization-wide digital transformation initiatives.

People also viewed

Renesas Electronics20 hours ago

Quality & Reliability Engineer – AI Enabling Power Modules

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Voltz20 hours ago

SRE Mid-level – Pipeline, Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Voltz20 hours ago

SRE Especialista – GCP

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Stone & Company21 hours ago

Site Reliability Engineer Manager

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
accesa.eu1 day ago

Senior DevOps Engineer

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Software Mind1 day ago

Senior Site Reliability Engineer – AI Experience Framework

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers