Remotery

Senior Site Reliability Engineer – Kubernetes

Posted 4 days ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Assist in the deployment, operation, and reliability of production services utilizing Kubernetes.

• Oversee service health and investigate incidents in production across distributed applications.

• Engage in on-call support, incident response, root cause analysis, postmortems, and enhancements in reliability.

• Diagnose issues related to application runtime, networking, and service-to-service interactions in collaboration with engineering teams.

• Facilitate CI/CD, GitOps-based deployments, observability, and monitoring of production environments.

• Operate within a client-directed backlog while adhering to established priorities.


⛳️ Requirements

• A minimum of 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a similar field.

• Recent, extensive hands-on experience in supporting Kubernetes-based production services.

• Preferred 3+ years of hands-on experience in production environments with Kubernetes.

• Proficiency in Kubernetes production operations, including deployment, scaling, rollout/rollback, resource tuning, and troubleshooting service-to-service interactions.

• Strong expertise in production incident response, including on-call duties, maintaining runbooks, conducting postmortems, and ensuring paging hygiene.

• Experience with Splunk for log aggregation, searching, and troubleshooting in production.

• Familiarity with Prometheus and Grafana, particularly in creating alert rules and dashboards.

• Knowledge of CI/CD and infrastructure-as-code practices for containerized deployments, including tools like Helm and GitOps solutions such as ArgoCD or Flux.

• Solid understanding of Linux and networking fundamentals, encompassing DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking.

• Experience in troubleshooting production issues across Node.js and JVM/Java services, with significant depth in at least one runtime environment.

• Skills in analyzing Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and investigating Java service latency.

• Familiarity with service-to-service authentication methods, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication.

• Additional skills in Web Components/Lit, server-side rendering or isomorphic runtime experience, canary rollout/multi-version production operations, distributed tracing and request-context correlation, KEDA/event-driven autoscaling, and enterprise platform integration layers are considered a plus.


🏝️ Benefits

• Competitive salary and a laptop provided.

• Opportunities for professional development and training.

• Work with advanced cloud and container technologies.

• Flexible work arrangements in a collaborative team environment.

• Contribute to organization-wide digital transformation initiatives.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers