
Mid-Level SRE
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Prevent production incidents by proactively identifying operational risks, potential points of failure, and excessive alerts before they escalate into issues.
• Engage in the development of the observability platform (logs, metrics, and tracing), helping to shape processes, SLIs, SLOs, and error budgets.
• Carry out root cause analyses and oversee postmortems, documenting insights gained and following up on action plans until they are fully executed.
• Tackle critical incidents by performing troubleshooting in both AWS and on-premises production settings.
• Discover FinOps opportunities and aid in enhancing cloud cost predictability.
• Work alongside the development team to continuously enhance application reliability and performance.
• Assist in scalability projects and the establishment of new infrastructure, emphasizing automation.
• Contribute to the advancement of the team’s SRE maturity through established practices, comprehensive documentation, and a culture of effective incident management.
• Extensive experience in troubleshooting production environments and managing distributed systems.
• Hands-on experience with AWS (EC2, networking, load balancing, IAM); familiarity with on-premises environments is advantageous.
• Proficient in Docker/Docker Compose, including the operation of containers in production settings.
• Experience with observability and monitoring tools (e.g., Grafana, Prometheus, Datadog, SigNoz, or similar) as well as OpenTelemetry (logs, metrics, and tracing).
• Solid understanding of SRE practices: SLI/SLO, error budgets, incident management and resolution, and conducting postmortems.
• Strong foundational knowledge of Linux, networking, and protocols (HTTP, TCP/IP, DNS).
• Excellent communication skills, independence, and resilience when operating during critical incidents, which may involve direct interaction with customers.
• Experience with automation tools (Python, Bash, Terraform, Ansible, or similar) is a plus.
• Familiarity with DevOps methodologies, including CI/CD and infrastructure as code (IaC), is a plus.
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Professional development opportunities and access to training resources.
• Comprehensive health and wellness benefits.
• Collaborative and inclusive company culture.
TEKsystems
Arctiq
GE Vernova
Ecosistemas
Get handpicked remote jobs straight to your inbox weekly.