Incident postmortem: the actions are the output, not the document
A postmortem that produces a beautifully written timeline and no completed actions has cost a team an afternoon and changed nothing.
Incident management ends when the service is back. Problem management is what nobody schedules — finding out why, and removing the cause.
Problem management is the discipline of identifying and removing the underlying causes of incidents, as distinct from resolving the incidents themselves. Incident management restores service as fast as possible, sometimes by a workaround that does not address why the failure happened. Problem management picks up from there.
The separation is deliberate and it is where most organisations quietly stop. Restoring service is urgent and visible; investigating a cause afterwards competes with planned work and loses. The result is familiar — the same category of incident recurring for years, each time resolved efficiently, never eliminated.
Problems are long-lived, cross-team and rarely urgent, which makes them exactly the kind of work that stalls. What keeps them moving is unglamorous: a named owner, a status that means something, and a recurring review where each open problem is either progressed, reprioritised or closed with a reason. Without that, the problem list becomes a graveyard that people stop adding to, because adding to it changes nothing.
Closing a problem because the incidents stopped is not the same as closing it because the cause was removed. Traffic patterns change, load moves, a dependency is retired — and the cause returns later with no record of the earlier investigation. Record what was actually done.
Incident volume is a poor measure of whether problem management is working, because it moves for many reasons. A better one is how often the same cause appears across incidents over time: a practice that works produces a falling number of repeat causes even when total incidents are flat. That measurement requires categorising incidents by cause rather than by symptom, which is a discipline in itself and worth the effort.
Ettex Board tracks problems and known errors with owners and status alongside other work, so they compete for attention honestly rather than sitting in a separate list; Ettex Records keeps the investigations and the workarounds where support can find them during the next incident; and the review that most often creates a problem record is covered in incident postmortem.
To be clear: this is task tracking and records, not an ITSM platform. What this covers is the practice — which is what fails, far more often than the tooling.
Identifying and removing the underlying causes of incidents, as distinct from restoring service during an incident.
A problem whose cause is understood but not yet resolved, recorded with a workaround so future incidents are handled faster.
Starting from data — trends, near misses, components with poor history — rather than waiting for incidents to recur.
By the frequency of repeat causes across incidents, not by incident volume, which moves for unrelated reasons.
A postmortem that produces a beautifully written timeline and no completed actions has cost a team an afternoon and changed nothing.
A snagging list records the defects found before handover. Its value is entirely in the timing — the same defect is a five-minute fix before completion and an argument afterwards.
Almost everything that goes wrong at a company event was decidable weeks earlier. A checklist ordered by deadline rather than by category is what turns the last week from firefighting into confirmation.