
Incident Operations Specialist
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of incident tooling operations. Ensure the reliability and configuration of incident.io, PagerDuty, Slack-based workflows, on-call rotations, and escalation paths. Monitor integrations and automations for problems and resolve any issues that arise.
• Develop and sustain AI-driven workflows. Design and implement repeatable automation for tasks such as incident thread summarization, postmortem draft creation, follow-up triage, severity classification, and data hygiene workflows. Transform one-time experiments into lasting systems that evolve over time, continuously improving them.
• Cultivate and support the IC community. Foster a community of practice for Incident Commanders and Support Leads. Conduct regular meetings, share insights from incidents, and mentor responders on best practices. Position yourself as a resource for others rather than just a system operator.
• Manage data and reporting systems. Develop, maintain, and troubleshoot dashboards and reports using tools like Databricks, Grafana, and Looker. Ensure data integrity, field completeness, and metric precision. Highlight trends and operational signals to the Incident Program Manager before they escalate into issues.
• Keep documentation and enablement resources up to date. Ensure playbooks, templates, and incident guides are current and effective under pressure. Identify areas where program-level guidance needs revision. Serve as the primary contact for inquiries regarding tooling and processes.
• Promote continuous improvement. Engage in incidents and postmortem evaluations. Recognize patterns in incidents and implement enhancements based on direct observations. Bring recurring challenges to the Incident Program Manager with actionable recommendations, not merely problems.
• Collaborate across the organization. Incidents involve Engineering, Support, GTM, Legal, and Finance. You facilitate coordination among these teams without needing the Program Manager to mediate every discussion. You understand who to involve, when to do so, and how to convey the necessary information for action.
• Support the incident program as it scales. As the program grows to encompass Support, Legal, PR, and Finance, assist in onboarding and enabling new stakeholder groups. Establish the infrastructure that allows the program to operate independently, even during your own PTO.
• Proven experience in incident response, technical operations, or a related reliability role.
• Proficient in incident tooling (incident.io, PagerDuty or similar) and Slack-based workflow automation.
• Comfortable working with SQL and reporting tools such as Databricks, Looker, and Grafana.
• Capable of diagnosing operational issues using logs, system context, and integration troubleshooting.
• Able to create lightweight automations, configure APIs, and prototype AI workflows independently, without needing an engineer.
• Actively utilizes AI in daily tasks for drafting, summarizing, building workflows, and triaging.
• Experienced in developing repeatable AI-driven workflows rather than one-time prompts.
• Applies verification and critical judgment to AI outputs, especially in high-pressure incident situations.
• Can articulate how their use of AI has enhanced speed, quality, or operational capacity.
• Proactively resolves issues when the direction is clear, even when the path requires further investigation.
• Exhibits strong attention to detail in configuration, data quality, and documentation.
• Maintains visibility through public status updates and proactive communication, eliminating the need for follow-ups.
• Demonstrates a history of respectfully disagreeing and committing, raising concerns early and executing once decisions are made.
• Communicates clearly in writing across various teams, including engineering, support, and operations.
• Proactively identifies and highlights risks, obstacles, and opportunities for improvement.
• Comfortable pushing back on off-program requests and articulating trade-offs.
• Capable of conveying highly technical incident information to non-technical audiences—such as Support teams, GTM partners, and leadership—clearly and without jargon.
• Provides equity opportunities.
• Offers bonuses.
Blink Health
Zscaler
Netwrix Corporation
NoGigiddy
Get handpicked remote jobs straight to your inbox weekly.