
Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Take ownership of the daily operations of a monitoring platform, which includes dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring.
• Actively analyze logs, error tracking, traces, and metrics to identify failures, regressions, and anomalies; triage, reproduce, and drive issues towards resolution.
• Adjust alert thresholds, monitoring logic, and notification routing to minimize noise and false positives while ensuring that customer-impacting issues are escalated to the appropriate personnel.
• Instrument services with valuable metrics, structured logs, and distributed traces.
• Collaborate with engineers to enhance code observability.
• Compose post-incident reviews, track remediation tasks, and incorporate lessons learned into monitors, runbooks, and system design.
• Define, assess, and report on service level indicators for essential customer-facing services.
• Enhance the reliability, scalability, and cost-effectiveness of cloud infrastructure, CI/CD pipelines, and release processes.
• Automate repetitive operational tasks.
• Maintain and refine runbooks, escalation paths, and operational documentation.
• Collaborate with enterprise InfoSec to address cybersecurity risks, including SSL/TLS cleanup, security headers, and DNS configuration hygiene.
• Monitor security findings from scans and audits until verified closure.
• Work alongside engineering, QA, product, and security teams to integrate reliability and observability into new features.
• Over 3 years of experience in site reliability engineering, DevOps, platform engineering, or a software engineering role focused on production.
• Practical experience in administering and building platforms such as New Relic, Grafana, Splunk, or Dynatrace, covering dashboards, monitors, log management, and APM.
• Strong skills in troubleshooting and root-cause analysis.
• Proficiency in at least one scripting or programming language, including Python, TypeScript/JavaScript, Go, or Bash.
• Ability to read application code to diagnose failures.
• Experience managing services in AWS, GCP, or Azure.
• A solid understanding of networking, containers, and Linux fundamentals.
• Familiarity with infrastructure as code tools such as Terraform, CloudFormation, or Pulumi.
• Acquainted with CI/CD tools like GitHub Actions.
• Experience with on-call duties, incident response, and post-incident review processes.
• Excellent written and verbal communication skills, capable of explaining reliability concerns and trade-offs to non-technical stakeholders.
• Health and dental insurance.
• Meal and food allowance.
• Childcare assistance.
• Extended paternity leave.
• Partnerships with gyms and health and wellness professionals through Wellhub (Gympass) TotalPass.
• Profit Sharing and Results Participation (PLR).
• Life insurance.
• Continuous learning platform (CI&T University).
• Discount club.
• Free online platform dedicated to physical, mental, and overall well-being.
• Pregnancy and responsible parenting course.
• Partnerships with online learning platforms.
• Language learning platform.
• Inclusion support and accommodations during the selection process.
• A dedicated Health and Well-being team, inclusion specialists, and affinity groups.
Entarian
BeyondTrust
BeyondTrust
Scribe
Get handpicked remote jobs straight to your inbox weekly.