
Site Reliability Engineer, SRE
Posted 8 hours ago

Posted 8 hours ago
This is a fully remote position, open to applicants in Poland, +2 more states.
• Take ownership of observability for essential product and user journeys within the squad.
• Establish, develop, and sustain relevant metrics, dashboards, and alerts.
• Define and uphold SLIs/SLOs for significant services and product-level metrics.
• Enhance monitoring, logging, tracing, and alerting across the squad’s systems.
• Serve as the primary responder for critical P0/P1 production incidents, including those outside of regular hours.
• Analyze production signals, pinpoint potential root causes, and begin addressing issues independently.
• Collaborate with other engineers when broader support or escalation is necessary.
• Engage in incident triage, mitigation, and postmortem reviews.
• Recognize recurring reliability challenges and drive enhancements to infrastructure, tooling, and incident-response procedures.
• Collaborate closely with backend and product engineers within a distributed, autonomous squad.
• Work directly with the client through GT’s Extended Team model and integrate deeply into the client’s team.
• Extensive prior experience as a Backend / Software Engineer, with senior-level practical expertise in Node.js and TypeScript.
• Hands-on experience in an SRE, Production Engineering, or a similar reliability-focused position.
• Strong production experience with AWS.
• Proficient in monitoring and observability across metrics, logging, tracing, and alerting.
• Practical experience in responding to production incidents, including triage, mitigation, and postmortems.
• Ability to interpret monitoring signals and independently investigate and start resolving production issues.
• Understanding of both application and infrastructure layers, rather than just infrastructure-only experience.
• Excellent communication skills and the ability to work independently within a distributed engineering team.
• Comfortable participating in out-of-hours incident response as part of the team’s coverage model.
• Nice to have: experience with Cloudflare and CloudWatch.
• Nice to have: familiarity with observability tools such as Sentry.
• Nice to have: experience in defining SLIs and SLOs for product-level metrics.
• Nice to have: experience with React Native or exposure to mobile application environments.
• Nice to have: previous experience with consumer mobile products or high-traffic B2C systems.
• Required English language proficiency (minimum selectable level specified in the application form).
• Must be willing to work on a B2B contract basis.
• Health insurance.
• Wellbeing budget.
• Sport coverage.
• Learning budget.
• 18 business days of paid vacation per year.
• Paid sick leave.
• All public holidays are paid days off.
• Encouraged trips to a client.
• Teambuilding and after-work activities.
CWILL
a37
Sigma Software Group
Applaudo
Get handpicked remote jobs straight to your inbox weekly.