Remotery

Site Reliability Engineer, SRE

Posted 8 hours ago

This is a fully remote position, open to applicants in Poland, +2 more states.

📋 Description

• Take ownership of observability for essential product and user journeys within the squad.

• Establish, develop, and sustain relevant metrics, dashboards, and alerts.

• Define and uphold SLIs/SLOs for significant services and product-level metrics.

• Enhance monitoring, logging, tracing, and alerting across the squad’s systems.

• Serve as the primary responder for critical P0/P1 production incidents, including those outside of regular hours.

• Analyze production signals, pinpoint potential root causes, and begin addressing issues independently.

• Collaborate with other engineers when broader support or escalation is necessary.

• Engage in incident triage, mitigation, and postmortem reviews.

• Recognize recurring reliability challenges and drive enhancements to infrastructure, tooling, and incident-response procedures.

• Collaborate closely with backend and product engineers within a distributed, autonomous squad.

• Work directly with the client through GT’s Extended Team model and integrate deeply into the client’s team.


⛳️ Requirements

• Extensive prior experience as a Backend / Software Engineer, with senior-level practical expertise in Node.js and TypeScript.

• Hands-on experience in an SRE, Production Engineering, or a similar reliability-focused position.

• Strong production experience with AWS.

• Proficient in monitoring and observability across metrics, logging, tracing, and alerting.

• Practical experience in responding to production incidents, including triage, mitigation, and postmortems.

• Ability to interpret monitoring signals and independently investigate and start resolving production issues.

• Understanding of both application and infrastructure layers, rather than just infrastructure-only experience.

• Excellent communication skills and the ability to work independently within a distributed engineering team.

• Comfortable participating in out-of-hours incident response as part of the team’s coverage model.

• Nice to have: experience with Cloudflare and CloudWatch.

• Nice to have: familiarity with observability tools such as Sentry.

• Nice to have: experience in defining SLIs and SLOs for product-level metrics.

• Nice to have: experience with React Native or exposure to mobile application environments.

• Nice to have: previous experience with consumer mobile products or high-traffic B2C systems.

• Required English language proficiency (minimum selectable level specified in the application form).

• Must be willing to work on a B2B contract basis.


🏝️ Benefits

• Health insurance.

• Wellbeing budget.

• Sport coverage.

• Learning budget.

• 18 business days of paid vacation per year.

• Paid sick leave.

• All public holidays are paid days off.

• Encouraged trips to a client.

• Teambuilding and after-work activities.

People also viewed

CWILL6 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a378 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group8 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo9 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch9 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job
Rimutee10 hours ago

DevOps / SRE, Part Time

Latin AmericaPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers