
SRE Engineer
Posted 8 hours ago

Posted 8 hours ago
This is a fully remote position, open to applicants in Brazil.
• Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability, monitoring, alerting, and Application Performance Management (APM).
• Track latency, traffic, errors, saturation, availability, and overall performance.
• Prevent, identify, and address incidents effectively.
• Lead root cause analysis and establish measures to prevent recurrence.
• Identify risks, bottlenecks, and single points of failure.
• Assist in designing resilient, scalable, and highly available solutions.
• Automate operational tasks to decrease manual intervention.
• Manage and enhance Kubernetes and Docker environments.
• Support capacity planning, business continuity, and disaster recovery initiatives.
• Participate in deployments and aid in application stabilization.
• Collaborate with teams to enhance reliability from the solution design phase.
• Develop and maintain dashboards, alerts, operational procedures, and documentation.
• Promote a culture of reliability, observability, and continuous improvement.
• Proven experience as a Site Reliability Engineer, SRE, or in a similar role.
• Practical experience with cloud platforms such as GCP, AWS, and/or Azure.
• Familiarity with Kubernetes and Docker.
• Experience with observability, monitoring, alerting, and APM tools.
• Understanding of SRE metrics and practices, including SLI, SLO, SLA, MTTR, and MTTD.
• Experience in managing, investigating, and resolving incidents.
• Knowledge in troubleshooting applications and infrastructure.
• Experience in administering Linux environments.
• Understanding of networking, security, performance, and high availability concepts.
• Experience with automation and Infrastructure as Code.
• Familiarity with CI/CD pipelines.
• Excellent communication skills and ability to collaborate with multidisciplinary teams.
• Analytical, proactive, collaborative, and prevention-focused mindset.
• Preferred: experience with GKE, EKS, or AKS.
• Preferred: knowledge of monitoring tools like Dynatrace, Datadog, Grafana, Prometheus, or similar.
• Preferred: experience with the ELK Stack, Elasticsearch, and Kibana.
• Preferred: knowledge of Terraform and Ansible.
• Preferred: experience with critical environments and distributed systems.
• Preferred: experience in financial institutions or regulated environments.
• Preferred: experience in capacity management and cloud cost optimization.
• Preferred: understanding of disaster recovery and business continuity.
• Preferred: experience in defining and managing error budgets.
• Preferred: certifications in Cloud, Kubernetes, or SRE.
• Meal voucher.
• Food allowance.
• Home office allowance.
• Medical insurance.
• Dental insurance.
• Life insurance.
• Birthday day off.
• TotalPass / Wellhub.
• Boon Saúde app.
• Discount partnerships.
• Discounts at partner businesses and educational institutions.
• Welcome kit.
• Onboarding program.
• Verity Learning.
• Verity Break.
• #VerityComVocê.
4Pharma Ltd
4Pharma Ltd
Virta Health
Segware
Get handpicked remote jobs straight to your inbox weekly.