
SRE Specialist
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Brazil.
• Engage in Site Reliability Engineering with a primary emphasis on observability.
• Collaborate with various technology teams to map, comprehend, and monitor the entire chain of services and systems.
• Lead technical evaluations and discovery workshops alongside Infrastructure, Networking, Database, Cloud, Security, Development, Architecture, Integration, Middleware, API, and Systems teams.
• Map the architecture and dependency chain of critical systems.
• Develop comprehensive Service Mapping, which includes infrastructure, applications, integrations, APIs, databases, queues, external services, and other dependencies.
• Identify integration pathways and evaluate the impact of components on service availability.
• Connect technical components to their corresponding business services.
• Detect monitoring and observability deficiencies.
• Establish and execute an observability strategy encompassing metrics, logs, traces, events, digital experience, and availability.
• Create service-oriented monitoring to oversee complete transaction and integration flows.
• Determine reliability metrics such as SLIs, SLOs, and service availability.
• Develop dashboards, alerts, correlations, and diagnostic tools.
• Minimize MTTD and MTTR while expediting root-cause analysis.
• Participate in corporate projects from inception, ensuring suitable observability requirements are met.
• Foster the growth of the observability culture and advocate for best practices among technical teams.
• Bachelor’s degree or an equivalent higher education qualification.
• Understanding of application architecture and distributed systems.
• Familiarity with server infrastructure and operating systems.
• Knowledge of networks and communication protocols.
• Proficiency in databases.
• Familiarity with REST APIs and system integrations.
• Understanding of microservices.
• Knowledge of containers and Kubernetes.
• Familiarity with cloud and hybrid environments.
• Understanding of queues and messaging systems.
• Knowledge of load balancers, proxies, and gateways.
• Familiarity with DNS, HTTP/HTTPS, TCP/IP, and TLS protocols.
• Knowledge of APM and Distributed Tracing methodologies.
• Understanding of centralized log management and analysis.
• Familiarity with infrastructure metrics and monitoring techniques.
• Knowledge of Service Mapping and Dependency Mapping.
• Understanding of SLI, SLO, SLA, and Error Budget concepts.
• Familiarity with Incident Management and root-cause analysis.
• Intermediate-level English proficiency is preferred.
• Desirable experience with tools such as Datadog, Elastic Stack/Elasticsearch/Logstash/Kibana/Elastic Agent, Sensedia/API Management, Zabbix, OpenTelemetry, Prometheus, and Grafana.
• A supportive environment for learning and professional development.
• Performance evaluations and feedback aimed at continuous improvement.
• Meal and/or food allowances.
• Medical and dental insurance coverage.
• Partnerships with pharmacies providing discounts on medications.
• Childcare assistance as per current policy.
• Collaboration with SESC, offering a range of cultural and leisure activities.
• Partnerships for language and technology training, along with access to a course platform.
• Payroll-deducted loans at favorable rates.
• Financial education initiatives.
• Corporate University and tailored learning paths.
• Referral Program with potential rewards and bonuses.
• Group life insurance coverage.
In All Media
Verity Group
Fingerprint
Endava
Get handpicked remote jobs straight to your inbox weekly.