
SRE Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability frameworks, monitoring systems, alerting mechanisms, and APM tools.
• Monitor key metrics including latency, traffic, errors, saturation, availability, and overall performance.
• Engage in incident prevention, identification, and resolution processes.
• Lead root cause analysis initiatives and establish actions to prevent future occurrences.
• Identify risks, bottlenecks, and potential single points of failure.
• Assist in designing resilient, scalable, and highly available solutions.
• Automate operational tasks to minimize manual efforts.
• Manage and enhance Kubernetes and Docker environments.
• Support strategies for capacity planning, business continuity, and disaster recovery.
• Participate in application deployments and provide support for stabilization efforts.
• Collaborate with teams to enhance reliability from the design phase of solutions.
• Create and maintain dashboards, alerts, procedures, and operational documentation.
• Foster a culture of reliability, observability, and continuous improvement.
• Proven experience as a Site Reliability Engineer, SRE, or in a similar role.
• Hands-on experience with cloud platforms like GCP, AWS, and/or Azure.
• Proficiency in Kubernetes and Docker.
• Experience with observability, monitoring, alerting, and APM processes.
• Familiarity with SRE metrics and practices including SLI, SLO, SLA, MTTR, and MTTD.
• Experience in managing, investigating, and resolving incidents.
• Knowledge of application and infrastructure troubleshooting techniques.
• Experience in administering Linux environments.
• Understanding of networking, security, performance, and high availability principles.
• Proficiency in automation and Infrastructure as Code methodologies.
• Experience with CI/CD pipelines.
• Strong communication skills, with the ability to collaborate with cross-functional teams.
• An analytical, proactive, collaborative, and prevention-oriented mindset.
• Nice to have: experience with GKE, EKS, or AKS.
• Nice to have: familiarity with Dynatrace, Datadog, Grafana, Prometheus, or comparable tools.
• Nice to have: experience with the ELK Stack, Elasticsearch, and Kibana.
• Nice to have: understanding of Terraform and Ansible.
• Nice to have: experience in mission-critical environments and distributed systems.
• Nice to have: experience within financial institutions or regulated environments.
• Nice to have: experience in capacity management and cloud cost optimization.
• Nice to have: knowledge of disaster recovery and business continuity practices.
• Nice to have: experience in defining and managing error budgets.
• Nice to have: certifications in Cloud, Kubernetes, or SRE.
• Meal voucher.
• Food allowance.
• Home office allowance.
• Health insurance.
• Dental insurance.
• Life insurance.
• Birthday Day Off.
• Total Pass / Wellhub app.
• Boon Saúde.
• Discount partnerships.
• Agreements with businesses and educational institutions.
• Welcome kit.
• Verity onboarding program.
• Verity Learning Interval.
• Great Place to Work certification and workplace improvement initiatives.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.