Senior Site Reliability Engineer, SRE

Posted Aug 20

This is a fully remote position, open to applicants in Brazil.

📋 Description

• Take ownership of observability for essential product and user journeys within the squad.

• Establish, develop, and sustain relevant metrics, dashboards, and alerts.

• Define and uphold SLIs/SLOs for crucial services and product-level metrics.

• Enhance monitoring, logging, tracing, and alerting across squad systems.

• Serve as the primary responder for critical P0/P1 production incidents, including those occurring after hours.

• Examine production signals, determine root causes, and begin addressing issues independently.

• Collaborate with engineers when broader support or escalation is necessary.

• Engage in incident triage, mitigation, and postmortem analyses.

• Recognize recurring reliability challenges and spearhead enhancements to infrastructure, tooling, and incident-response processes.

• Collaborate closely with backend and product engineers in a distributed, autonomous squad.


⛳️ Requirements

• Significant prior experience as a Backend / Software Engineer.

• Senior-level, hands-on expertise in Node.js and TypeScript.

• Practical experience in an SRE, Production Engineering, or a similar reliability-focused position.

• Strong production experience with AWS.

• Familiarity with Cloudflare and CloudWatch.

• Experience in monitoring and observability across metrics, logging, tracing, and alerting.

• Hands-on experience responding to production incidents, including triage, mitigation, and postmortem evaluations.

• Ability to interpret monitoring signals and independently investigate and start resolving production issues.

• Understanding of both application and infrastructure layers.

• Excellent communication skills and the ability to work autonomously within a distributed engineering team.

• Comfortable participating in out-of-hours incident responses.

• Required proficiency in English (application form requires candidates to select their level).


🏝️ Benefits

• Health insurance.

• Wellbeing budget.

• Sport coverage.

• Learning budget.

• 18 business days of paid vacation each year.

• Paid sick leaves.

• All public holidays are paid days off.

• Direct management by the client.

• Access to the same tools and technologies as the client.

• Opportunities for trips to the client.

• Team-building and after-work activities.

• Long-term projects lasting several years.

People also viewed

FCamara Consulting & Training4 hours ago

Senior SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sequoia Connect4 hours ago

DevOps Engineer, Java, Cloud

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FundCount6 hours ago

DevOps Team Lead

US flagUnited States OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
GE Vernova6 hours ago

Senior Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152.4k – $254k/year
ApplyView job
NASCO6 hours ago

Delivery DevOps Agile Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IBM8 hours ago

Senior DevOps Engineer, Systems

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers