What ServiceNow Major Incident Management Actually Looks Like Inside a Real War Room

September 28, 2026 The ServiceNow Guy 10 min read
What ServiceNow Major Incident Management Actually Looks Like Inside a Real War Room

A retail CIO called me on a Sunday night. Their point-of-sale integration had been down for four hours across 380 stores. Their MIM process, on paper, was textbook. Dedicated channel in Teams. Incident commander named in the runbook. Bridge open. Status page updating every 15 minutes. Communications template pre-approved by comms and legal.

None of it was working. The bridge had 34 people on it, most of them muted, most of them unclear on why they had been invited. The incident commander was a service desk manager who had been told two months earlier that this role was hers on paper. She had never actually run one. The status page updates were being drafted by a comms person who was pulling status from the bridge, which was pulling status from a screenshare of a Kibana dashboard that only one engineer could read. And the ServiceNow major incident record itself had been created 90 minutes into the outage, after someone remembered it should exist.

This is what most ServiceNow major incident management looks like when the wheels come off. Not because the tooling is bad. The tooling is fine. Because the process was designed for the audit, not for the crisis.

The gap between the runbook and the room

Every enterprise I have worked with in the last decade has a major incident management process document. Most of them are 30 to 60 pages long. They cover escalation criteria, communication cadences, RACI matrices, post-incident review requirements, and executive notification thresholds. They are reviewed annually. They pass audits.

Almost none of them survive contact with a real Priority 1.

The reason is not the document. The reason is that the document was written by people who have never had to run a P1 at 3am while three vice presidents are asking for updates and the engineering team is arguing about which of two hypotheses to test first. The document assumes the room is quiet. Real rooms are not quiet.

A working ServiceNow major incident management setup has to answer four questions in the first ten minutes of an incident, every single time, without anyone having to look anything up. Who is running this. Where is the bridge. What is the current hypothesis. What is the next update time. Everything else is a nice-to-have.

The incident commander role that actually works

The most common failure mode I see is treating “incident commander” as a title assigned in a spreadsheet rather than a rehearsed role. The person named as incident commander in your ServiceNow major incident process needs three things: authority, availability, and reps.

Authority means the CIO or head of infrastructure has told the organisation, in writing, that when this person is running an incident, their decisions stand. If a director wants to override them, that override goes through the CIO, not through arguing with the IC on the bridge. Without that authority the IC becomes a note-taker.

Availability means they are actually reachable in the shift they cover. Not “usually reachable”. Reachable. If your named IC for European weekends is on holiday, the process needs to route to someone who has been briefed and knows they are the IC that weekend. This is a rostering problem, not a documentation problem.

Reps means they have actually run incidents. Tabletop exercises count for something, but the muscle memory that lets an IC keep 20 engineers focused on a single hypothesis for 45 minutes without wandering off comes from having done it live, probably badly, several times. If your MIM process has three named ICs and only one of them has ever actually run a P1, you have one IC and two names on a page.

The ServiceNow side of this is straightforward. The major incident record should have an assigned Incident Commander field that is separate from Assignment Group and Assigned To. It should be populated within the first five minutes of the record being created. The bridge details, the current status, and the next update time should all live in fields on that record, updated by the IC or a scribe, not lost in Teams chat history.

The bridge, the scribe, and the status board

The 34-person bridge I described at the start is the second most common failure mode. Bridges balloon because everyone thinks they need to be on them. Directors want visibility. Vendors are pulled in because someone thought their input might help. Product owners join because they saw a Slack notification.

A working war room bridge has three tiers, and the incident commander enforces the boundary between them.

The core bridge is the people actively diagnosing and fixing. Usually four to eight engineers plus the IC and a scribe. Cameras optional, mics unmuted only when speaking, no side conversations.

The extended bridge is a separate channel for people who need situational awareness but are not part of the fix. Directors, product owners, second-line vendor contacts. They read the status board and the scribe’s timeline. They do not talk into the core bridge.

The comms bridge is a third channel where the comms lead, the customer success lead, and the executive notifier draft external messaging using the scribe’s timeline as source material. They do not interrupt the core bridge with drafting questions.

The scribe role is the one most organisations skip and then regret. The scribe sits on the core bridge and writes a running timeline directly into the ServiceNow major incident record, ideally into a dedicated Work Notes stream that filters out chatter. Every hypothesis tested, every action taken, every result observed. In real time, not from memory afterwards. This is what makes the post-incident review actually useful and what lets the comms bridge write updates without pulling engineers off the diagnosis.

If your ServiceNow setup does not have a clean way for a scribe to log timeline entries without noise from bots, integrations, and automated updates, fix that before your next P1. A cluttered timeline is worse than no timeline because it forces the reviewer to reconstruct events under time pressure. This is exactly the kind of thing our ServiceNow consulting services get called in for when a MIM programme is technically live but operationally broken.

Escalation criteria that trigger, not gather dust

Every MIM process has escalation criteria. Most of them are written in a way that requires human interpretation to trigger, which means they get triggered late, or after the fact, or not at all.

“Customer-facing impact of significant severity” is not a trigger. It is a topic for debate at 4am. Working criteria look like this: any incident where P1 has been declared, or where more than one business service is degraded, or where the resolution ETA exceeds 60 minutes from declaration, automatically escalates to major incident status and pages the IC on call.

Automation matters here. If your ServiceNow instance can detect a P1 declaration and immediately create the major incident record, populate the IC field from an on-call rota, open the bridge, and notify the extended channel, you have removed 15 minutes of coordination overhead from the first hour of every P1. That is the difference between a well-run incident and a chaotic one.

The rota integration is where most instances fall down. If your IC on-call schedule lives in a Google Sheet that the service desk consults manually, your automation is broken. It needs to live in a system ServiceNow can query at incident-creation time, whether that is a proper on-call tool like PagerDuty or a ServiceNow On-Call Scheduling configuration. Either works. The Google Sheet does not.

Post-incident review as the actual product

The output of a ServiceNow major incident management programme is not the incident record. It is the review that comes out of it, and specifically the actions that the review generates.

Most PIR templates I see are structured around the “five whys” and the “what went well / what went badly” format. Both are fine as far as they go. Neither of them, by themselves, produce action.

What produces action is the discipline of turning every finding into a specific owned change with a due date, tracked in ServiceNow, reviewed at the next PIR to see whether it was actually done. If your MIM programme has been running for a year and there is no list of changes that were made because of PIR findings, the programme is not working. It is theatre.

The changes do not need to be dramatic. Most of them are small. A monitoring alert that should have fired earlier. A runbook step that was ambiguous. A dashboard that took too long to load. A vendor contact that turned out to be wrong. The point is not the size of the change. The point is that the incident produced a specific improvement that would have prevented or shortened the next occurrence.

This is also where the mim discipline ties directly back to reducing mean time to resolution across the board. A programme that captures ten small, specific improvements per quarter compounds. Twelve months later the MTTR curve has bent noticeably. This does not happen because someone did a big transformation project. It happens because every P1 fed one or two improvements back into the system.

Where to start, practically

If you are looking at your current major incident management setup and recognising some of these gaps, four moves have the highest return.

Get your incident commander roster right first. Named ICs, real authority, actual availability, at least two of them with live reps. Everything else depends on this.

Second, split your bridges. One core, one extended, one comms. Publish the rule and enforce it. The IC has permission to remove people from the core bridge who are not adding value to the fix.

Third, get the scribe role staffed and the timeline logging clean. If your ServiceNow Work Notes on major incidents are full of integration noise, filter it. Make the timeline the single source of truth for the whole organisation during the incident.

Fourth, wire your escalation criteria to trigger automatically. P1 declaration to major incident creation to IC paging to bridge open should be one workflow, not four manual steps.

If you want a structured read of where your current MIM setup stands, our 10-Day Instance Health Report covers major incident readiness as one of the six dimensions we audit. It will tell you, specifically, where your process breaks under load and what to fix first. Most of what we find is not exotic. It is the same handful of gaps I described above, showing up in different combinations.

Mladen Milic runs Milic Media Kft, a boutique ServiceNow consultancy delivering implementation, health audits and HRSD work across the EU. Reach him at mladen@milicmedia.com.

Need Help With ServiceNow?

Our consulting team can help you implement what you've read about — and more.

Leave a Reply

Your email address will not be published. Required fields are marked *