Remotery

Senior Site Reliability Engineer – AI Experience Framework

Posted 19 hours ago

This is a fully remote position, open to applicants in Poland.

📋 Description

• Take ownership of Kubernetes deployment and ensure operational health for Framework services, which includes scaling, rollout/rollback strategies, and resource optimization.

• Develop and uphold production observability using Grafana dashboards and Prometheus alerting rules across the SSR runtime and Glide platform layer.

• Identify and resolve Node.js production incidents, addressing issues such as event-loop stalls, heap growth, V8 isolate exhaustion, and isolate-pool scheduling challenges.

• Troubleshoot and resolve JVM production incidents, focusing on GC pressure, thread dumps, and platform-service latency.

• Manage incident response through runbooks, on-call rotations, postmortems, and ensure paging hygiene.

• Lead CI/CD initiatives and infrastructure-as-code practices for Kubernetes manifests/Helm and deployment pipelines.

• Collaborate with framework engineering to pinpoint reliability gaps via capacity planning, load testing, and chaos/failure injection techniques.

• Advocate for production reliability considerations during architecture reviews for new framework functionalities.


⛳️ Requirements

• Experience in production operations/SRE, specifically with practical Kubernetes deployment, scaling, and incident response.

• Direct operational experience in troubleshooting Node.js in production, including heap snapshots, CPU profiles, event-loop blocking, and process/worker isolation models.

• Proven experience in troubleshooting JVM-based services in a production environment, including working with GC logs, thread dumps, and JVM tuning.

• Hands-on experience in building and maintaining Prometheus alerting rules and Grafana dashboards from the ground up.

• Strong understanding of Linux and networking fundamentals, including DNS, load balancing, TCP/HTTP semantics, HTTP/2, and Kubernetes service-to-service networking.

• Familiarity with CI/CD and infrastructure-as-code practices for containerized deployments, including Helm and GitOps tools like ArgoCD/Flux or their equivalents.

• Demonstrated experience managing on-call rotations, creating runbooks, and leading postmortems.

• Experience with Splunk for log aggregation, searching, and troubleshooting in production environments.

• Practical production experience with Valkey/Redis, focusing on key design, TTL/eviction tuning, and tenant-scoped cache invalidation.

• Knowledge of managing service-to-service mTLS, including certificate issuance and rotation, format conversion, and JWT-based service authentication.

• Proficient spoken and written English skills.

• Familiarity with server-side rendering architectures and isomorphic-runtime failure modes.

• Experience implementing multi-version/canary rollout strategies.

• Working knowledge of the ServiceNow Glide platform or a similar enterprise platform integration layer.

• Experience with distributed tracing and request-context correlation.

• Familiarity with event-driven autoscaling solutions such as KEDA ScaledObjects driven by PromQL triggers.


🏝️ Benefits

• Flexible employment options and remote work opportunities.

• Engage in international projects with leading global clients.

• Opportunities for international business travel.

• A non-corporate work atmosphere.

• Language classes available.

• Access to internal and external training.

• Private healthcare and insurance coverage.

• Multisport card benefits.

• Initiatives focused on well-being.

People also viewed

accesa.eu19 hours ago

Senior DevOps Engineer

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF19 hours ago

DevSecOps Engineer, Mid Level – Top Secret

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$89.6k – $184.4k/year
ApplyView job
Sumsub21 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
SailPoint22 hours ago

Senior DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Union Home Mortgage Corp.22 hours ago

Infrastructure DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Point Wild (Formerly Pango Group)23 hours ago

Lead DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers