
Senior Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Brazil.
• Manage and enhance environments after the Go-Live phase
• Oversee the availability, performance, and overall health of production services
• Develop and improve observability practices
• Ensure the platform is comprehensible, accessible, and recoverable
• Extensive experience as a Site Reliability Engineer (SRE), DevOps Engineer, or a related position
• Demonstrated experience in maintaining highly available production environments
• In-depth knowledge of AWS
• Practical experience with monitoring, observability, and alerting tools
• Understanding of metrics, logs, and distributed tracing
• Familiarity with incident management and response processes
• Knowledge in capacity planning, performance, availability, and resilience
• Experience in automating operational workflows
• Awareness of SRE principles, including SLI, SLO, and SLA
• Competence in troubleshooting and conducting root cause analyses
• Experience with applications and services that handle real-time traffic
• Hands-on, analytical, and adept at problem-solving
• Proficient in English at an advanced level
• Competitive salary and compensation package
• Opportunities for professional growth and development
• Flexible working hours and remote work options
• Access to health and wellness programs
Slate Auto
Funding Xchange
Leidos
LeoLabs
Get handpicked remote jobs straight to your inbox weekly.