
Senior Site Reliability Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Brazil.
• Ensure the accessibility, scalability, and performance of applications and services.
• Implement and enhance observability practices, including metrics, logs, and traces.
• Create and maintain dashboards for monitoring system health indicators.
• Define and manage alerting systems, emphasizing effective alerts and reducing noise.
• Identify and resolve incidents while conducting root cause analysis (RCA).
• Collaborate closely with development teams to promote continuous improvement through DevOps practices.
• Automate operational routines and monitoring processes.
• Assist in defining and tracking SLIs, SLOs, and SLAs.
• Contribute to fostering a culture of reliability and resilience engineering.
• Experience in Identity and Access Management (IAM).
• Knowledge of Ping Identity solutions is a significant advantage.
• Familiarity with observability, with experience using platforms like Datadog, Elastic, and Grafana.
• Strong communication skills, with the capability to track requests and collaborate with multidisciplinary teams.
• Open to individuals with disabilities (PwD).
• Participation in the selection process does not require any payment of participation or hiring fees.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.