
Senior Site Reliability Engineer – AI Experience Framework
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Poland.
• Take ownership of Kubernetes deployment and ensure operational health for Framework services, which includes scaling, rollout/rollback strategies, and resource optimization.
• Develop and uphold production observability using Grafana dashboards and Prometheus alerting rules across the SSR runtime and Glide platform layer.
• Identify and resolve Node.js production incidents, addressing issues such as event-loop stalls, heap growth, V8 isolate exhaustion, and isolate-pool scheduling challenges.
• Troubleshoot and resolve JVM production incidents, focusing on GC pressure, thread dumps, and platform-service latency.
• Manage incident response through runbooks, on-call rotations, postmortems, and ensure paging hygiene.
• Lead CI/CD initiatives and infrastructure-as-code practices for Kubernetes manifests/Helm and deployment pipelines.
• Collaborate with framework engineering to pinpoint reliability gaps via capacity planning, load testing, and chaos/failure injection techniques.
• Advocate for production reliability considerations during architecture reviews for new framework functionalities.
• Experience in production operations/SRE, specifically with practical Kubernetes deployment, scaling, and incident response.
• Direct operational experience in troubleshooting Node.js in production, including heap snapshots, CPU profiles, event-loop blocking, and process/worker isolation models.
• Proven experience in troubleshooting JVM-based services in a production environment, including working with GC logs, thread dumps, and JVM tuning.
• Hands-on experience in building and maintaining Prometheus alerting rules and Grafana dashboards from the ground up.
• Strong understanding of Linux and networking fundamentals, including DNS, load balancing, TCP/HTTP semantics, HTTP/2, and Kubernetes service-to-service networking.
• Familiarity with CI/CD and infrastructure-as-code practices for containerized deployments, including Helm and GitOps tools like ArgoCD/Flux or their equivalents.
• Demonstrated experience managing on-call rotations, creating runbooks, and leading postmortems.
• Experience with Splunk for log aggregation, searching, and troubleshooting in production environments.
• Practical production experience with Valkey/Redis, focusing on key design, TTL/eviction tuning, and tenant-scoped cache invalidation.
• Knowledge of managing service-to-service mTLS, including certificate issuance and rotation, format conversion, and JWT-based service authentication.
• Proficient spoken and written English skills.
• Familiarity with server-side rendering architectures and isomorphic-runtime failure modes.
• Experience implementing multi-version/canary rollout strategies.
• Working knowledge of the ServiceNow Glide platform or a similar enterprise platform integration layer.
• Experience with distributed tracing and request-context correlation.
• Familiarity with event-driven autoscaling solutions such as KEDA ScaledObjects driven by PromQL triggers.
• Flexible employment options and remote work opportunities.
• Engage in international projects with leading global clients.
• Opportunities for international business travel.
• A non-corporate work atmosphere.
• Language classes available.
• Access to internal and external training.
• Private healthcare and insurance coverage.
• Multisport card benefits.
• Initiatives focused on well-being.
accesa.eu
ICF
Sumsub
SailPoint
Get handpicked remote jobs straight to your inbox weekly.