
Technical Enterprise Incident Manager
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Lead and facilitate incident bridge calls that encompass infrastructure, application, network, cloud, security, and vendor teams.
• Accelerate service restoration while ensuring accurate timelines, effective communications, and executive updates.
• Assess incidents based on business impact and operational risk to prioritize response efforts.
• Oversee escalation procedures and involve leadership as necessary.
• Track SLA compliance and incident response metrics to ensure performance standards are met.
• Conduct Post-Incident Reviews and ensure thorough documentation of Root Cause Analysis and corrective actions.
• Promote automation for incident detection, triage, and response workflows.
• Maintain organized stakeholder communications during significant incidents.
• Validate runbooks, service dependency maps, and technical documentation for accuracy.
• Utilize log aggregation and APM tools for effective troubleshooting and incident analysis.
• Collaborate with application teams to address issues and implement root-cause solutions.
• Develop and enhance capabilities for monitoring, alerting, and dashboarding.
• Analyze trends, KPIs, and operational metrics to pinpoint reliability risks.
• Support strategies for resiliency, including redundancy, failover, capacity planning, and performance optimization.
• Leverage Datadog for monitoring, alert correlation, dashboards, incident investigation, and performance assessment.
• Participate in an after-hours on-call incident management rotation.
• Create and maintain incident management procedures, runbooks, and knowledge articles.
• Ensure precise ticket documentation in ServiceNow.
• Promote ongoing service improvement in line with ITIL and SRE best practices.
• Work collaboratively with cross-functional teams to enhance communication, escalation processes, and workflows.
• Support audit, compliance, and operational reporting obligations.
• Must be a U.S. citizen.
• Capability to obtain and maintain the necessary Public Trust level clearance.
• Bachelor's Degree with 5 years of experience, or a High School diploma or equivalent combined with 9 years of experience.
• 5+ years of experience in Cloud Incident Management, Operations Engineering, NOC, SRE, or Application/Production Support environments.
• Proven experience leading enterprise Major Incident response efforts in a 24/7 operational setting.
• Strong comprehension of ITIL Incident and Problem Management processes.
• 3+ years of experience working with AWS cloud services.
• Familiarity with monitoring and observability platforms like Datadog, Cloudcraft, or similar tools.
• Proficient in using ServiceNow or comparable ITSM platforms.
• Strong analytical, troubleshooting, and organizational skills.
• Excellent written and verbal communication abilities, capable of facilitating meetings and briefing technical teams as well as executive leadership.
• Preferred: Experience in a Site Reliability Engineering (SRE) or Cloud Platform DevOps environment.
• Preferred: Experience supporting federal, healthcare, financial, or other highly regulated sectors.
• Preferred: Hands-on experience with Windows/Linux Servers, networking concepts, cloud platforms (AWS, Azure, or GCP), load balancers, proxies, DNS, and firewalls.
• Potential eligibility for overtime.
• Shift differential may be available.
• Discretionary bonus may be available.
Neogen Corporation
Clear Star
Get handpicked remote jobs straight to your inbox weekly.