
Senior Monitoring Analyst
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Brazil.
• Oversee the availability, performance, and capacity of assets, links, and essential services.
• Lead the analysis of alarms, events, metrics, and indicators during high-impact or complex incidents.
• Conduct diagnostics and technical escalations, identifying causes, impacts, priorities, and containment measures.
• Perform advanced troubleshooting related to connectivity, infrastructure, collection, and monitoring services.
• Design and maintain hosts, templates, items, preprocessing, triggers, macros, discovery rules (LLD), actions, and event correlation within Zabbix.
• Develop and enhance dashboards, variables, datasources, transformations, and alerts using Grafana.
• Investigate collection failures through SNMP, ICMP, agents, Syslog, APIs, proxies, and various other integrations.
• Define and review thresholds, dependencies, severities, and suppression rules to improve alarm accuracy.
• Analyze capacity and availability trends to anticipate risks, saturation, and potential degradations.
• Conduct root cause analyses and recommend corrective and preventive measures.
• Ensure logging, escalations, and technical communications align with established SLAs and workflows.
• Mentor analysts, perform technical reviews, and disseminate knowledge.
• Establish standards, document solutions, and contribute to the advancement of monitoring and NOC operations.
• Extensive experience with Zabbix in production settings.
• In-depth knowledge of distributed architecture, proxies, templates, items, preprocessing, LLD (low-level discovery), triggers, dependencies, macros, actions, and event correlation.
• Advanced proficiency with Grafana, including datasources, variables, transformations, alerts, and dashboards.
• Familiarity with SNMPv2c/v3, ICMP, agents, Syslog, APIs, and webhooks.
• Comprehensive understanding of TCP/IP, addressing, DNS, VLANs, switching, routing, latency, and packet loss.
• Ability to interpret metrics, establish baselines, and correlate events and trends effectively.
• Proficient with tools such as ping, traceroute, MTR, dig, curl, and tcpdump.
• Knowledge of Linux for analyzing services, logs, processes, resources, and integrations.
• Experience in incident and problem management, root cause analysis, escalation, SLAs, and operating at scale.
• Strong background in monitoring, infrastructure, or NOC operations, particularly in managing critical incidents and high-availability environments.
• Hands-on expertise in Zabbix and Grafana, encompassing configuration, troubleshooting, and the evolution of production environments.
• Ongoing development opportunities.
• Technical support resources.
• Incentives for relevant training and certifications.
CVS Health
IVC Evidensia UK
DIGESTIVE HEALTH
Advanced Cooling Technologies, Inc.
Get handpicked remote jobs straight to your inbox weekly.