
SRE Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Establish and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability, monitoring, alerting, and APM solutions.
• Track latency, traffic, errors, saturation, availability, and performance metrics.
• Prevent, identify, and resolve incidents effectively.
• Lead root cause analysis efforts and define actions to avoid recurrence.
• Identify risks, bottlenecks, and single points of failure in systems.
• Assist in designing resilient, scalable, and highly available solutions.
• Automate operational tasks to minimize manual intervention.
• Manage and enhance Kubernetes and Docker environments.
• Support capacity planning, business continuity, and disaster recovery initiatives.
• Participate in deployments and aid in application stabilization efforts.
• Collaborate with teams to enhance reliability starting from the solution design phase.
• Develop and maintain dashboards, alerts, procedures, and operational documentation.
• Foster a culture of reliability, observability, and continuous improvement.
• Proven experience as a Site Reliability Engineer, SRE, or in a similar capacity.
• Hands-on expertise with cloud platforms including GCP, AWS, and/or Azure.
• Proficiency in Kubernetes and Docker technologies.
• Experience with observability, monitoring, alerting, and APM tools.
• Familiarity with SRE metrics and practices such as SLI, SLO, SLA, MTTR, and MTTD.
• Experience in managing, investigating, and resolving incidents.
• Knowledge of application and infrastructure troubleshooting techniques.
• Experience in administering Linux environments.
• Understanding of networking, security, performance, and high availability principles.
• Experience in automation and Infrastructure as Code methodologies.
• Familiarity with CI/CD pipeline methodologies.
• Excellent communication skills and ability to collaborate with cross-functional teams.
• Possess an analytical, proactive, collaborative, and prevention-focused mindset.
• Experience with GKE, EKS, or AKS.
• Knowledge of tools such as Dynatrace, Datadog, Grafana, Prometheus, or similar.
• Familiarity with the ELK Stack, Elasticsearch, and Kibana.
• Understanding of Terraform and Ansible.
• Experience in mission-critical environments and distributed systems.
• Background in financial institutions or regulated environments.
• Experience with capacity management and cloud cost optimization practices.
• Knowledge of disaster recovery and business continuity strategies.
• Experience in defining and managing error budgets.
• Relevant cloud, Kubernetes, or SRE certifications.
• Availability for employment under the CLT regime, as specified in the application form.
• Meal allowance.
• Food allowance.
• Home office allowance.
• Health insurance.
• Dental insurance.
• Life insurance.
• Birthday day off.
• TotalPass / Wellhub membership.
• Boon Saúde app access.
• Discount partnerships.
• Collaborations with businesses and educational institutions.
• Welcome kit for new employees.
• Comprehensive onboarding process.
• Verity Learning resources.
• Verity Break initiatives.
• #VerityComVocê community engagement.
• Great Place to Work certification.
Entarian
BeyondTrust
BeyondTrust
Scribe
Get handpicked remote jobs straight to your inbox weekly.