
Associate Service Delivery Manager
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in India.
• Oversee the complete incident lifecycle, which includes detection, triage, escalation, and resolution in accordance with SLA commitments.
• Collaborate with Cloud Operations, SRE, Engineering, and Support teams during incidents.
• Assess incident severity and establish escalation pathways based on impact, urgency, and system criticality.
• Deliver structured and timely incident updates to internal stakeholders and customer-facing teams.
• Assist in root cause analysis by gathering inputs, validating findings, and monitoring corrective and preventive actions to completion.
• Recognize recurring issues and participate in problem management initiatives.
• Track incident trends, response times, and SLA compliance while recommending data-driven enhancements.
• Develop and refine runbooks, playbooks, and standard operating procedures.
• Enhance service resilience and operational processes in collaboration with Engineering, SRE, and Cloud Operations.
• Utilize AI-assisted analysis for incident triage, signal interpretation, and root cause investigations.
• Leverage AI-driven insights to enhance prioritization, minimize MTTD and MTTR, and reduce customer impact.
• Identify and execute automation opportunities in incident management, RCA, reporting, and operational workflows.
• Employ observability, telemetry, monitoring signals, and AI analytics to detect anomalies, risks, and service degradation.
• Contribute to AIOps capabilities including anomaly detection, event correlation, alert noise reduction, and predictive risk identification.
• Optimize alerting mechanisms, dashboards, and operational runbooks.
• Evaluate the impact of AI/AIOps through reductions in incidents, the quality of alerts, and the adoption of automation.
• Proven experience in service delivery, incident management, or service operations within a high-availability environment.
• Practical experience in managing incidents and coordinating cross-functional teams in time-sensitive situations.
• Strong analytical and problem-solving capabilities.
• Familiarity with incident lifecycle management, RCA processes, and problem management practices.
• Knowledge of AI tools such as Copilot or similar.
• Exposure to AIOps concepts such as observability, telemetry, anomaly detection, and intelligent alerting.
• Ability to interpret telemetry, logs, and monitoring signals.
• Experience with ITSM tools like Jira, ServiceNow, PagerDuty, or comparable platforms.
• Excellent communication skills for engaging with both technical and non-technical stakeholders.
• Capacity to handle multiple priorities while maintaining focus on service restoration and operational results.
• Preferred: experience in SaaS or cloud-based environments and distributed systems.
• Preferred: familiarity with monitoring platforms like Datadog, New Relic, or similar.
• Preferred: exposure to automation, scripting, or workflow optimization.
• Preferred: understanding of AWS or Azure.
• Preferred: experience in contributing to AIOps-enabled initiatives.
• Preferred: ITIL certification or equivalent service management training.
• Bachelor's degree in a related field or equivalent practical experience.
• Ability to maintain the confidentiality, integrity, and availability of information assets.
• Responsibility for employee and customer data privacy and completion of required privacy training.
• Remote-first work environment.
• Employee Resource Groups.
• Coffee with Mark sessions with the CEO.
• Microsoft Teams communities focused on wellness, art, pets, family, and parenting.
• Special guest sessions on issues affecting employees.
Signature Aviation
Allstate
Zayo Group
97th Floor
Get handpicked remote jobs straight to your inbox weekly.