Major Incident Management

Major incident war-room patterns and command roles.

What ServiceNow Major Incident Management Actually Looks Like Inside a Real War Room

What ServiceNow Major Incident Management Actually Looks Like Inside a Real War Room

A retail CIO called me on a Sunday night. Their point-of-sale integration had been down for four hours across 380 stores. Their MIM process, on paper, was textbook. Dedicated channel in Teams. Incident commander named in the runbook. Bridge open. Status page updating every 15 minutes. Communications template pre-approved by comms and legal. None of it was working. The bridge had 34 people on it, most of them muted, most of them unclear on why they had been invited. The incident commander was a service desk manager who had been told two months earlier that this role was hers on paper. She had never actually run one. The status page updates were being drafted by a comms person who was pulling status from the bridge, which was pulling status from a screenshare of a Kibana dashboard that only one engineer could read. And the ServiceNow major incident record itself had been created 90 minutes into the outage, after someone remembered it should exist. This is what most ServiceNow major incident management looks like when the wheels come off. Not because the tooling is bad. The tooling is fine. Because the process was designed for the audit, not for the crisis.

ServiceNow Major Incident Management: The War-Room Patterns That Actually Hold Under Pressure

ServiceNow Major Incident Management: The War-Room Patterns That Actually Hold Under Pressure

A head of IT operations at a European retailer called me on a Saturday night last quarter. Their checkout flow had been intermittent for ninety minutes, the store was three days from a public earnings update, and the war-room bridge had eighteen people on it. Six of them were vendors. Two were lawyers. Nobody was running the call. The incident commander had been pulled into a separate executive bridge to brief the CFO and had not returned. The bridge had drifted into a roundtable of vendors describing what their own monitoring did and did not show. The retailer's ServiceNow major incident management process had a beautiful page in the runbook describing the bridge etiquette. None of it was happening. This is the part of ServiceNow major incident management that nobody puts in the demo. The platform handles the workflow plumbing well. It opens the major incident record, it pages the commander, it sets the comms cadence, it logs the timeline. What it cannot do, and what nine in ten implementations get wrong, is the operating model around the workflow. The war room is a human process running on top of a technical record. If the human process is not designed, the technical record...