
Senior Site Reliability Engineer – Kubernetes
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Poland.
• Assist with the deployment, operation, and reliability of production services utilizing Kubernetes.
• Oversee service health and investigate incidents in production across distributed applications.
• Engage in on-call support, incident response, root cause analysis, postmortems, and enhancements in reliability.
• Diagnose issues related to application runtime, networking, and service-to-service communication in collaboration with engineering teams.
• Facilitate CI/CD, GitOps-based deployments, observability, and monitoring of production environments.
• Operate within a client-directed backlog and adhere to established priorities.
• Take ownership of production reliability for the AI Experience Framework stack from start to finish, encompassing Kubernetes, observability, and addressing issues in Node.js and JVM systems.
• Minimum of 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a related field.
• Significant recent hands-on experience in supporting production services based on Kubernetes.
• Preferred 3+ years of direct experience with production Kubernetes.
• Expertise in Kubernetes production operations, which includes deployment, scaling, rollout/rollback, resource tuning, and service-to-service troubleshooting.
• Strong background in production incident response, which includes on-call duties, runbooks, postmortems, and maintaining paging hygiene.
• Familiarity with Splunk for log aggregation, searching, and troubleshooting in production environments.
• Experience with Prometheus and Grafana for creating alert rules and dashboards.
• Proficiency in CI/CD and infrastructure-as-code practices for containerized deployments, including usage of Helm and GitOps tools such as ArgoCD or Flux.
• Strong fundamentals in Linux and networking, covering DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking.
• Experience in troubleshooting production issues across Node.js and JVM/Java services, with a strong focus on at least one runtime environment.
• Knowledge of service-to-service authentication methods, including mTLS, certificate rotation, format conversion, and JWT-based authentication.
• Excellent spoken and written English skills.
• Additional skills that are beneficial include Web Components/Lit experience, server-side rendering or isomorphic runtime experience, canary rollout/multi-version production operations, distributed tracing, KEDA or event-driven autoscaling, and enterprise platform integration experience.
• Flexible employment options and remote working arrangements.
• Opportunities to work on international projects with leading global clients.
• Potential for international business travel.
• A non-corporate work environment.
• Language classes offered.
• Access to internal and external training programs.
• Private healthcare and insurance coverage.
• Multisport card available.
• Initiatives focused on employee well-being.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.