
Problem Manager – ITIL, ServiceNow, Root Cause Analysis
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Oversee enterprise-wide initiatives for Problem Management, Root Cause Analysis, and Incident Prevention.
• Conduct 5 Whys RCA sessions for P1 incidents and recurring P2 incidents.
• Analyze logs, monitoring tools, dashboards, and change records to determine timelines and contributing factors.
• Utilize AI-assisted tools for investigations, timeline creation, RCA documentation, and knowledge capture.
• Record root causes and corrective actions in ServiceNow Problem Management records.
• Lead the Problem Management lifecycle from investigation to closure.
• Collaborate with technical teams to develop Remediation Action Plans.
• Monitor remediation progress, escalate risks, and ensure stakeholder accountability.
• Validate solutions and assess service reliability, operational risk, and improvements in incident recurrence.
• Identify recurring trends across infrastructure, cloud, application, and network incidents.
• Suggest enhancements in monitoring, automation, and architecture.
• Facilitate weekly Problem Review meetings.
• Maintain executive dashboards and manage Problem Management reporting.
• Provide executive-ready summaries, status updates, and RCA reports for customers and senior leadership.
• Enhance governance processes, operational standards, runbooks, and reporting practices.
• Promote automation and AI-enabled workflows.
• A minimum of 7 years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related fields.
• Strong knowledge of ITIL Problem Management, Incident Management, and operational governance.
• Proven experience conducting Root Cause Analysis, 5 Whys investigations, and post-incident reviews.
• Practical experience with ServiceNow or similar ITSM platforms.
• Exceptional analytical, documentation, facilitation, and stakeholder management skills.
• Ability to influence cross-functional teams and foster accountability without direct authority.
• Experience in creating executive-level communications, dashboards, and operational reports.
• Comprehensive technical understanding of Windows and Linux platforms.
• Familiarity with VMware and Hyper-V virtualization technologies.
• Skills in performance troubleshooting across CPU, memory, storage, and I/O.
• Knowledge of SAN and NAS environments, including storage performance, redundancy, and resiliency.
• Understanding of networking concepts such as routing, switching, VLANs, DNS, firewalls, load balancing, and packet-flow analysis.
• Ability to interpret network latency and packet loss.
• Knowledge of AWS and Microsoft Azure, focusing on cloud networking, compute, identity, and storage services.
• Understanding of cloud architecture and failure-mode analysis.
• Familiarity with APIs, microservices, distributed systems, CI/CD, and deployment pipelines.
• Awareness of application defects, configuration drift, and dependencies as potential causes of operational incidents.
• ITIL Foundation certification or higher is preferred.
• Experience in supporting large-scale enterprise environments is preferred.
• Knowledge of operational analytics, trend analysis, and KPI reporting is preferred.
• Exposure to automation, AI-enabled workflows, or AIOps practices is preferred.
• Experience with executive stakeholders and client-facing incident communications is preferred.
• Eligibility for bonuses or incentives based on business needs.
• Health insurance coverage.
• Optional dental and vision plans.
• Life and disability insurance.
• Retirement savings plan.
• Paid holidays.
• Paid time off (PTO), vacation, and/or sick leave.
Jeevan Technologies (a Nobl Q company)
QTS Data Centers
Get handpicked remote jobs straight to your inbox weekly.