
SRE Engineer
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in Brazil.
• Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability, monitoring, alerting, and APM solutions.
• Oversee latency, traffic, errors, saturation, availability, and overall performance.
• Prevent, identify, and address incidents effectively.
• Lead root cause analysis efforts and establish actions to avoid future occurrences.
• Identify risks, bottlenecks, and single points of failure within systems.
• Assist in designing resilient, scalable, and highly available solutions.
• Automate operational processes to minimize manual tasks.
• Manage and enhance Kubernetes and Docker environments.
• Aid in capacity planning, business continuity, and disaster recovery initiatives.
• Engage in deployments and support application stabilization efforts.
• Collaborate with teams to enhance reliability from the initial design phase onward.
• Develop and maintain dashboards, alerts, procedures, and operational documentation.
• Foster a culture of reliability, observability, and continuous improvement.
• Proven experience as a Site Reliability Engineer, SRE, or in a similar role.
• Practical experience with cloud environments, including GCP, AWS, and/or Azure.
• Familiarity with Kubernetes and Docker technology.
• Experience with observability, monitoring, alerting, and APM tools.
• Understanding of SRE metrics and methodologies, such as SLI, SLO, SLA, MTTR, and MTTD.
• Experience in managing, investigating, and resolving incidents.
• Knowledge of troubleshooting applications and infrastructure.
• Experience administering Linux systems.
• Understanding of networking, security, performance, and high availability principles.
• Proficiency in automation and Infrastructure as Code practices.
• Familiarity with CI/CD pipeline processes.
• Strong communication skills with the ability to collaborate with multidisciplinary teams.
• Analytical, proactive, collaborative, and focused on prevention.
• Preferred: Experience with GKE, EKS, or AKS.
• Preferred: Familiarity with tools such as Dynatrace, Datadog, Grafana, Prometheus, or similar.
• Preferred: Experience with the ELK Stack, Elasticsearch, and Kibana.
• Preferred: Knowledge of Terraform and Ansible.
• Preferred: Experience in mission-critical environments and distributed systems.
• Preferred: Experience within financial institutions or regulated environments.
• Preferred: Background in capacity management and cloud cost optimization.
• Preferred: Knowledge of disaster recovery and business continuity protocols.
• Preferred: Experience in defining and managing error budgets.
• Preferred: Relevant cloud, Kubernetes, or SRE certifications.
• Meal voucher
• Food allowance
• Home office allowance
• Medical insurance
• Dental insurance
• Life insurance
• Birthday day off
• TotalPass / Wellhub app
• Boon Saúde
• Discount partnerships
• Agreements with businesses and educational institutions
• Welcome kit
• Verity onboarding
• Verity Learning Interval
• Great Place to Work certification
In All Media
Verity Group
Fingerprint
Endava
Get handpicked remote jobs straight to your inbox weekly.