
Associate Director, Observability and Service Reliability
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Canada.
• Take ownership of and enhance the strategy for enterprise observability and service reliability.
• Establish enterprise-wide standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and essential technology services.
• Set expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and visibility of business transactions.
• Detect monitoring deficiencies, redundant features, excessive alerts, and areas for enhancing visibility.
• Transition the organization towards proactive, predictive, and automated operations.
• Develop and advance the service reliability framework.
• Collaborate with technical service owners to outline monitoring requirements and measures for service reliability.
• Act as the enterprise technical authority and provide architectural advice on observability, monitoring, and service reliability.
• Create monitoring patterns, reference architectures, standards, and reusable capabilities.
• Assess emerging technologies in observability, AIOps, automation, analytics, and service reliability.
• Oversee the strategy, architecture, governance, adoption, and optimization of Dynatrace, Nexthink, and related monitoring platforms.
• Manage critical technology and vendor partnerships.
• Promote the adoption of Dynatrace across applications, infrastructure, cloud, and digital services.
• Lead the Nexthink and Digital Employee Experience strategy, focusing on proactive remediation and automation.
• Enhance event management, signal quality, event correlation, anomaly detection, automated diagnostics, and proactive remediation.
• Integrate observability platforms with ITSM, incident management, automation, collaboration, configuration, and operational data platforms.
• Establish governance for monitoring standards, tools, integrations, data quality, licensing, and adoption.
• Create executive and operational reports on service health, reliability, performance, and user experience.
• Lead and nurture professionals in observability, monitoring, reliability, and platform engineering.
• Build relationships across application, platform, infrastructure, cloud, network, cybersecurity, Digital Workplace, DevOps, SRE, Service Management, and business teams.
• Collaborate with Incident, Problem, Change, Major Incident Management, and Operational Resilience teams to enhance detection, recovery, and prevention.
• A Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field, or equivalent professional experience.
• A minimum of 8 years of experience in observability, application performance, service reliability, infrastructure, cloud, engineering, or enterprise technology operations.
• At least 3 years of experience in a technical or people leadership role.
• Proven experience in designing and managing monitoring and observability capabilities within complex enterprise settings.
• Experience with Dynatrace and familiarity with Nexthink or similar Digital Employee Experience technologies.
• Knowledge of observability technologies, monitoring architectures, application performance management, infrastructure monitoring, event management, logging, tracing, synthetic monitoring, and user experience monitoring.
• Understanding of modern enterprise applications, cloud technologies, containers, APIs, networks, databases, infrastructure, and distributed architectures.
• Experience in defining service health, reliability measures, Service Level Indicators, Service Level Objectives, and operational performance metrics.
• Ability to influence senior technical leaders and convey complex technical concepts in terms of business risk and operational outcomes.
• Strong leadership, architectural, analytical, problem-solving, communication, and stakeholder management skills.
• Preferred: Advanced experience with Dynatrace in a large enterprise environment.
• Preferred: Experience in implementing or scaling Nexthink.
• Preferred: Familiarity with ServiceNow and enterprise event management platforms.
• Preferred: Experience in Site Reliability Engineering, DevOps, AIOps, automation, OpenTelemetry, cloud-native monitoring, and contemporary observability architectures.
• Preferred: Experience in establishing enterprise monitoring standards or observability reference architectures.
• Preferred: Background in complex, global, or highly regulated enterprise environments.
• Flexible and supportive work environment.
• Prioritization of employee well-being.
• A hybrid-friendly company culture.
• Be Well programs aimed at supporting financial, mental, physical, and social health.
• Personalized development objectives.
• Continuous feedback mechanisms.
• Certification opportunities with Microsoft, Google, and Amazon.
• Coaching and practical learning experiences.
• Access to cutting-edge learning opportunities.
• Tools for career advancement and in-demand skill access.
Alzheimer's Association®
The College Board
The College Board
Cargill
Get handpicked remote jobs straight to your inbox weekly.