← All postsHow-to

Problem management: the work that happens after service is restored

Incident management ends when the service is back. Problem management is what nobody schedules — finding out why, and removing the cause.

How-toP

Problem management is the discipline of identifying and removing the underlying causes of incidents, as distinct from resolving the incidents themselves. Incident management restores service as fast as possible, sometimes by a workaround that does not address why the failure happened. Problem management picks up from there.

The separation is deliberate and it is where most organisations quietly stop. Restoring service is urgent and visible; investigating a cause afterwards competes with planned work and loses. The result is familiar — the same category of incident recurring for years, each time resolved efficiently, never eliminated.

Reactive and proactive problem management

  • Reactive problem management starts from incidents that have happened: recurring failures, high-impact single events, and clusters that share a cause.
  • Proactive problem management starts from data: error rates trending upwards, near misses, capacity approaching limits, components with a poor history.
  • A known error is a problem whose cause is understood but not yet fixed, recorded with its workaround so the next incident is resolved faster.
  • The known error record is the highest-value artefact of the whole practice, and the one most often kept in someone’s head.

Give problems owners and a review, or they decay

Problems are long-lived, cross-team and rarely urgent, which makes them exactly the kind of work that stalls. What keeps them moving is unglamorous: a named owner, a status that means something, and a recurring review where each open problem is either progressed, reprioritised or closed with a reason. Without that, the problem list becomes a graveyard that people stop adding to, because adding to it changes nothing.

Closing a problem because the incidents stopped is not the same as closing it because the cause was removed. Traffic patterns change, load moves, a dependency is retired — and the cause returns later with no record of the earlier investigation. Record what was actually done.

Count causes, not tickets

Incident volume is a poor measure of whether problem management is working, because it moves for many reasons. A better one is how often the same cause appears across incidents over time: a practice that works produces a falling number of repeat causes even when total incidents are flat. That measurement requires categorising incidents by cause rather than by symptom, which is a discipline in itself and worth the effort.

Where the records live

Ettex Board tracks problems and known errors with owners and status alongside other work, so they compete for attention honestly rather than sitting in a separate list; Ettex Records keeps the investigations and the workarounds where support can find them during the next incident; and the review that most often creates a problem record is covered in incident postmortem.

To be clear: this is task tracking and records, not an ITSM platform. What this covers is the practice — which is what fails, far more often than the tooling.

Frequently asked

What is problem management?

Identifying and removing the underlying causes of incidents, as distinct from restoring service during an incident.

What is a known error?

A problem whose cause is understood but not yet resolved, recorded with a workaround so future incidents are handled faster.

What is proactive problem management?

Starting from data — trends, near misses, components with poor history — rather than waiting for incidents to recur.

How should it be measured?

By the frequency of repeat causes across incidents, not by incident volume, which moves for unrelated reasons.

SL
Written by Sofia L.

Part of the Ettex team — writing about product, engineering and the future of work.

More posts
Get the best of the Ettex blogProduct news, guides and tips — straight to your inbox, no spam.