
Observability Engineer
Posted Sep 16

Posted Sep 16
This is a fully remote position, open to applicants in Brazil.
• Design and oversee a cohesive observability framework that encompasses metrics, logs, traces, and events.
• Spearhead First-of-a-Kind implementations by assessing new observability technologies and transforming them into secure, repeatable, and production-ready patterns.
• Establish enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
• Define and manage Service Level Indicators, Objectives, and error budgets.
• Act as the senior technical escalation point during critical P1/P2 incident war rooms and perform evidence-based Root Cause Analysis.
• Minimize MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping.
• Promote Observability-as-Code and infrastructure automation.
• Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.
• Integrate observability platforms with ITSM, Netcool, and CI/CD pipelines.
• Design observability solutions for Docker, Kubernetes, microservices, and multi-cloud environments.
• Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
• Ensure telemetry pipelines are secure by design.
• Lead Knowledge Transfer programs, vendor transitions, and operational readiness handovers for global 24x7 teams.
• Mentor cross-functional engineering teams and shape enterprise technology roadmaps.
• Proficiency with IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, and InfluxDB.
• Experience with SolarWinds, Netcool, Elastic/Splunk.
• Knowledge of Kubernetes, Docker, OpenShift, and cloud platforms like AWS/Azure/GCP.
• Familiarity with Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
• Skills in Ansible, Terraform, Python, Bash, Webhooks, and CI/CD (GitHub Actions/GitLab/Jenkins).
• Understanding of ServiceNow, ITIL 4, and advanced Major Incident Management.
• 12+ years of overall IT experience.
• At least 5 to 7 years in a Lead Architect, SRE, or Principal Observability Engineer role within a large enterprise setting.
• Demonstrated success in transitioning organizations from legacy monitoring systems to proactive, automated observability platforms.
• Hands-on experience in developing scalable, secure telemetry pipelines and time-series databases.
• Extensive background in leading FOAK rollouts and complex vendor/operations transition (KT) initiatives.
• Preferred certifications include CKA, Cloud Architect (AWS/Azure), or specific APM/Observability vendor certifications.
• CLT hiring - 40 hours weekly workload.
• Remote work opportunity.
• Health Insurance - SulAmérica Prestige (available for legal dependents).
• Dental Insurance - SulAmérica (available for legal dependents).
• Life Insurance - Prudential (24x salary).
• Private Pension - Metlife (up to 6% company match).
• Meal Voucher and Internet allowance - Flash (R$1.000,00 per month).
• Employee Assistance Program.
• Wellness program.
• Individual performance compensation program, based on eligibility.
• Equity grant under the Associate Equity Appreciation Program, contingent on eligibility.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.