SRE Engineer

Posted 1 day ago

This is a fully remote position, open to applicants in Brazil.

📋 Description

• Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.

• Implement observability frameworks, monitoring systems, alerting mechanisms, and APM tools.

• Monitor key metrics including latency, traffic, errors, saturation, availability, and overall performance.

• Engage in incident prevention, identification, and resolution processes.

• Lead root cause analysis initiatives and establish actions to prevent future occurrences.

• Identify risks, bottlenecks, and potential single points of failure.

• Assist in designing resilient, scalable, and highly available solutions.

• Automate operational tasks to minimize manual efforts.

• Manage and enhance Kubernetes and Docker environments.

• Support strategies for capacity planning, business continuity, and disaster recovery.

• Participate in application deployments and provide support for stabilization efforts.

• Collaborate with teams to enhance reliability from the design phase of solutions.

• Create and maintain dashboards, alerts, procedures, and operational documentation.

• Foster a culture of reliability, observability, and continuous improvement.


⛳️ Requirements

• Proven experience as a Site Reliability Engineer, SRE, or in a similar role.

• Hands-on experience with cloud platforms like GCP, AWS, and/or Azure.

• Proficiency in Kubernetes and Docker.

• Experience with observability, monitoring, alerting, and APM processes.

• Familiarity with SRE metrics and practices including SLI, SLO, SLA, MTTR, and MTTD.

• Experience in managing, investigating, and resolving incidents.

• Knowledge of application and infrastructure troubleshooting techniques.

• Experience in administering Linux environments.

• Understanding of networking, security, performance, and high availability principles.

• Proficiency in automation and Infrastructure as Code methodologies.

• Experience with CI/CD pipelines.

• Strong communication skills, with the ability to collaborate with cross-functional teams.

• An analytical, proactive, collaborative, and prevention-oriented mindset.

• Nice to have: experience with GKE, EKS, or AKS.

• Nice to have: familiarity with Dynatrace, Datadog, Grafana, Prometheus, or comparable tools.

• Nice to have: experience with the ELK Stack, Elasticsearch, and Kibana.

• Nice to have: understanding of Terraform and Ansible.

• Nice to have: experience in mission-critical environments and distributed systems.

• Nice to have: experience within financial institutions or regulated environments.

• Nice to have: experience in capacity management and cloud cost optimization.

• Nice to have: knowledge of disaster recovery and business continuity practices.

• Nice to have: experience in defining and managing error budgets.

• Nice to have: certifications in Cloud, Kubernetes, or SRE.


🏝️ Benefits

• Meal voucher.

• Food allowance.

• Home office allowance.

• Health insurance.

• Dental insurance.

• Life insurance.

• Birthday Day Off.

• Total Pass / Wellhub app.

• Boon Saúde.

• Discount partnerships.

• Agreements with businesses and educational institutions.

• Welcome kit.

• Verity onboarding program.

• Verity Learning Interval.

• Great Place to Work certification and workplace improvement initiatives.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF13 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam17 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers