
SRE Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in Brazil.
• Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability, monitoring, alerting, and application performance management (APM).
• Track latency, traffic, errors, saturation, availability, and performance metrics.
• Prevent, identify, and address incidents effectively.
• Conduct root cause analyses and establish measures to prevent future occurrences.
• Identify risks, bottlenecks, and single points of failure within systems.
• Assist in designing resilient, scalable, and highly available solutions.
• Automate operational tasks to minimize manual interventions.
• Manage and enhance Kubernetes and Docker environments.
• Support capacity planning, business continuity, and disaster recovery initiatives.
• Participate in deployments and assist with application stabilization.
• Collaborate with teams to enhance reliability from the solution design phase.
• Develop and maintain dashboards, alerts, procedures, and operational documentation.
• Foster a culture focused on reliability, observability, and continuous improvement.
• Experience as a Site Reliability Engineer (SRE) or in a similar role.
• Practical experience with cloud platforms such as GCP, AWS, and/or Azure.
• Proficiency in Kubernetes and Docker technologies.
• Experience with observability, monitoring, alerting, and APM tools.
• Understanding of SRE metrics and practices, including SLI, SLO, SLA, MTTR, and MTTD.
• Experience in managing, investigating, and resolving incidents.
• Familiarity with application and infrastructure troubleshooting.
• Experience in administering Linux environments.
• Knowledge of networking, security, performance, and high availability concepts.
• Experience with automation and Infrastructure as Code methodologies.
• Familiarity with CI/CD pipeline implementations.
• Excellent communication skills and the capability to collaborate with cross-functional teams.
• Analytical, proactive, collaborative, and focused on prevention.
• Experience with GKE, EKS, or AKS.
• Knowledge of monitoring tools such as Dynatrace, Datadog, Grafana, Prometheus, or similar.
• Experience with the ELK stack, including Elasticsearch and Kibana.
• Familiarity with Terraform and Ansible.
• Experience with critical environments and distributed systems.
• Previous experience in financial institutions or regulated environments.
• Experience in capacity management and optimizing cloud costs.
• Knowledge of disaster recovery and business continuity planning.
• Experience in defining and managing error budgets.
• Relevant certifications in Cloud, Kubernetes, or SRE.
• Meal allowance.
• Food allowance.
• Home office allowance.
• Health insurance.
• Dental insurance.
• Life insurance.
• Birthday day off.
• Total Pass / Wellhub membership.
• Boon Saúde app access.
• Discount partnerships.
• Discounts at partner establishments and educational institutions.
• Welcome kit.
• Comprehensive onboarding process.
• Access to Verity Learning resources.
• Participation in Verity Break initiatives.
• Engage with #VerityComVocê community.
• Opportunities for professional development courses.
• Work in a Great Place to Work-certified environment.
In All Media
Fingerprint
Endava
CVS Health
Get handpicked remote jobs straight to your inbox weekly.