Root cause analysis: getting past the first plausible answer
Root cause analysis works out why something went wrong deeply enough that fixing it prevents a recurrence. Its main difficulty is stopping too early, at an answer that sounds satisfying.
An incident response plan exists so that the first hour is not improvised. Roles, severity levels and notification deadlines are the parts you cannot work out while it is happening.
An incident response plan sets out what happens when something goes wrong with your systems or data: who leads, how severity is judged, what is done first, who must be told and by when. Its purpose is narrow and important — to remove decisions from the worst possible moment to make them, which is while the incident is running and everyone is anxious.
The plan does not need to cover every kind of incident. Most small companies face a short list: an account compromised, ransomware or malware, data sent to the wrong recipient, a supplier breached, an outage of something customers depend on. Write for those.
Find out your notification deadlines before you need them. Many data-protection regimes require notifying a regulator within a short fixed window of becoming aware of a personal-data breach, and contracts often impose their own, shorter obligations to notify customers. These are jurisdiction- and contract-specific — establish yours in advance and write them into the plan, because working them out during an incident is how deadlines get missed.
Someone should be writing down what happened and when, from the first minute: what was observed, what was decided, who was told, what was changed. It matters for three reasons — regulators and insurers ask for a timeline, the post-incident review is worthless without one, and memory reconstructs events wrongly under stress. It is a boring job that nobody volunteers for, which is why the plan should name the role rather than hope someone starts.
In Ettex, the plan itself lives in Ettex Docs — two pages with version history, threaded comments while it is agreed, and a share link so it is reachable rather than buried. The incident log is a table in Ettex Records: timestamp, observation, decision, who was notified, all with revision history so the record itself is auditable. Contacts sit in Ettex Contacts and should be exported to a copy outside your systems, holding statements are drafted in advance in Ettex Docs, and Ettex Chat carries the internal coordination as long as the incident does not affect it — which is exactly why the contact list must also exist offline.
The boundary, plainly: Ettex is not a security product. There is no monitoring, detection, alerting or paging, no SIEM, no forensics, no on-call rota, and nothing here will tell you an incident is happening. It holds the plan, the log and the contacts. Detection and technical response come from your IT provider or security tooling — and if your plan depends on Ettex being available, keep an offline copy.
A short document defining what counts as an incident, how severity is judged, who leads and communicates, the first containment actions, notification obligations, and how the incident is closed and reviewed.
About two pages plus a contact list. Anything longer will not be read during an incident, which is the only time it matters.
It depends on your jurisdiction and your contracts — several data-protection regimes set a short fixed window from becoming aware, and customer contracts often impose tighter ones. Establish your specific deadlines in advance and write them into the plan.
One named role with a deputy, authorised to make decisions and spend money. Separate that from whoever handles communication, since doing both badly is the usual outcome.
Because regulators and insurers ask for one, the post-incident review depends on it, and memory under stress reconstructs events inaccurately. Name someone to record it from the first minute.
Causes and changes, not blame. Hold it within a week while detail is fresh, and track the resulting actions like any other work — otherwise the same incident recurs.
An incident response plan is two pages that remove decisions from the worst moment: who leads, what severity means, what to do in the first hour, who to tell and by when. Write it calmly, keep it reachable offline, and rehearse it once a year.
Root cause analysis works out why something went wrong deeply enough that fixing it prevents a recurrence. Its main difficulty is stopping too early, at an answer that sounds satisfying.
A code of conduct says how people here are expected to behave and what happens when they do not. Its value is not aspiration — it is having decided the hard cases before one arrives.
A project charter names the objective, the boundaries, the sponsor and the person authorised to run the work. Its main use comes months later, when people disagree about what was agreed.