
Incident Operations Lead – EMEA/AMER
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in United States.
• Oversee the team managing Alpaca's most significant incidents.
• Develop the incident operations function, which includes the severity model, escalation and communication pathways, 24x7 follow-the-sun coverage, and improvement KPIs.
• Recruit and certify Incident Commanders while establishing rotations across APAC, EMEA, and AMER with regional transitions.
• Hold a rostered incident-command position and mentor the team during live incidents.
• Facilitate game days, tabletop exercises, and simulations.
• Foster a culture of blameless reviews.
• Enhance severity maturity in collaboration with Risk regarding financial and regulatory significance.
• Manage escalation paths, unanswered-page protocols, engineering-leadership thresholds, and ownership of the service catalogue.
• Link engineers and technical support with partner communications teams during incidents.
• Ensure precise update frequencies and conduct retrospectives with SRE.
• Create post-incident packages that include ticketed, assigned, and tagged actions.
• Share clear incident insights with the engineering organization.
• Take ownership of response and mitigation KPIs, including definitions, data hygiene, baselines, and targets.
• Report overdue postmortem reviews along with subsequent actions.
• Develop a documented, version-controlled, and deployable incident-command product.
• Manage the roadmap for AI workflows and agents supporting incident setup, timelines, RCAs, action packages, and follow-up reminders.
• You have established an incident command or major-incident function, taking responsibility for the severity model, creating the roster, and promoting adoption across teams.
• Over 5 years of experience in production engineering, SRE, or technical operations, including direct management of high-severity incidents.
• Proven ability to lead a distributed team across various time zones and manage a 24x7 rotation.
• Capacity to influence engineers you do not directly manage and justify severity decisions.
• Experience in building reliable metrics and differentiating metric improvement from genuine enhancements.
• Strong scope discipline during outages.
• Excellent written and verbal communication skills, including delivering executive briefings during incidents.
• Knowledge of FinTech and the trust implications of API-driven financial platforms.
• Utilization of AI and agentic automation to reduce manual tasks.
• Formal incident command training, such as ITIL, Major Incident Management, or crisis management, is a plus.
• Experience with certification, game day, or drill programs is a plus.
• Familiarity with modern incident management and on-call platforms is a plus.
• Experience in developing and maintaining service catalogues or ownership registries is a plus.
• Experience in converting incident follow-ups into funded roadmap projects is a plus.
• Knowledge of DORA, Reg SCI, FINRA, or similar reporting obligations is a plus.
• Experience in online securities trading, capital markets, or other regulated, market-sensitive sectors is a plus.
• Competitive Salary & Stock Options
• Health Benefits
• New Hire Home-Office Setup: One-time USD $500
• Monthly Stipend: USD $150 per month via a Brex Card
CVS Health
Mondelēz International
Humana
Pipeline Medical
Get handpicked remote jobs straight to your inbox weekly.