
Senior Site Reliability Engineer – Kubernetes
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in Canada.
• Assist in the deployment, operation, and maintenance of production services utilizing Kubernetes.
• Oversee service health and investigate incidents occurring in production across distributed applications.
• Engage in on-call support, incident response, root cause analysis, postmortem reviews, and enhancements in reliability.
• Diagnose runtime issues, networking challenges, and service-to-service problems in collaboration with engineering teams.
• Facilitate CI/CD, GitOps-based deployments, observability, and monitoring of production environments.
• Operate within a client-directed backlog and adhere to set priorities.
• A minimum of 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a similar role.
• Recent and substantial hands-on experience in supporting production services based on Kubernetes.
• Preferred: At least 3 years of hands-on experience with Kubernetes in production.
• Familiarity with Kubernetes production operations, including deployment, scaling, rollout/rollback, resource optimization, and troubleshooting service-to-service issues.
• Extensive experience in production incident response, covering on-call duties, runbooks, postmortems, and paging protocols.
• Proficiency in Splunk for log aggregation, search capabilities, and troubleshooting in production environments.
• Experience with Prometheus and Grafana, particularly in creating alerting rules and dashboards.
• Knowledge of CI/CD and infrastructure-as-code practices for containerized deployments, including tools like Helm and GitOps platforms such as ArgoCD or Flux.
• Solid understanding of Linux and networking fundamentals, covering DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking.
• Expertise in Node.js production troubleshooting, including heap snapshots, CPU profiling, event-loop blocking, memory management, worker/process isolation, and V8 isolates or similar runtime models.
• Proficiency in JVM/Java production troubleshooting, including GC log analysis, thread dump examination, JVM tuning, and investigating service latency in Java applications.
• Familiarity with in-memory caching systems like Redis/Valkey, including key design, TTL/eviction tuning, and cache invalidation strategies.
• Experience with service-to-service authentication mechanisms, such as mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication.
• Experience with Web Components/Lit is a plus.
• Knowledge of server-side rendering or isomorphic runtime environments is an advantage.
• Familiarity with canary rollout and multi-version production operations is desirable.
• Experience with distributed tracing and request-context correlation is beneficial.
• Knowledge of KEDA or event-driven autoscaling is a plus.
• Familiarity with enterprise platform integration layers is advantageous.
• Competitive salary and laptop provided.
• Opportunities for professional development and training.
• Engage with state-of-the-art cloud and container technologies.
• Flexible work arrangements and a collaborative team environment.
• Contribute to organization-wide digital transformation initiatives.
Renesas Electronics
Voltz
Voltz
Stone & Company
Get handpicked remote jobs straight to your inbox weekly.