Root Cause Analysis That Changes Something
Why most investigations stop one level too early, how to tell a real root cause from a convenient one, and the step almost nobody performs that turns a risk register into a measure of whether the programme works.
Root cause analysis is the structured search for the condition that, if corrected, would prevent an event from recurring — as distinct from the immediate cause, which is simply what happened last before things went wrong.
The distinction is the whole discipline, and it is where most investigations quietly fail. An analysis concluding that an operator made a mistake has identified an immediate cause and stopped. The corrective action that follows is retraining, which is comfortable, cheap and almost always ineffective — because the next operator, working under the same conditions, will make the same mistake.
This guide covers how to run a root cause analysis that reaches a cause worth correcting, why the five whys technique fails when applied casually, and the measurement step almost nobody performs.
Human error is where investigations stop, not where they should
There is a reliable warning sign that a root cause analysis stopped early: the conclusion names a person. Somebody failed to follow the procedure, somebody forgot a step, somebody took a shortcut. All of those may be factually accurate, and none of them is a root cause.
The useful next question is why a competent person, doing their job, behaved that way. Usually the answer is that the correct action was slower, harder or less obvious than the incorrect one — and that is a condition you can change, unlike a person's attention on a particular Tuesday.
This matters commercially as well as ethically. An organisation whose investigations routinely conclude in individual fault will find that reporting dries up, because people learn that raising an event produces an inquiry into who is to blame. The investigation quality and the reporting rate collapse together, and the second failure is usually noticed long after the first.
The five whys, done properly
The technique is simple enough to be dismissed and useful enough to keep. Its value is that it forces the chain to be written down, so the point where the reasoning becomes weak is visible to everyone rather than concealed inside a narrative paragraph. A worked example, drawn from a familiar failure:
Notice where the chain lands. Stopping at why three produces a corrective action of tightening the belt, which fixes this compressor and nothing else. Stopping at why four produces a conversation with the technician, which is the human-error trap. Reaching why five produces a change to the checklist — which protects every compressor on the site, including the ones nobody has looked at yet.
Two cautions about the technique. It is not a rule that there are exactly five levels; some chains resolve in three and some need eight, and padding to reach five produces nonsense. And it produces a single line of causation, which most real events do not have — a fishbone diagram covering method, machine, material, environment, measurement and people is better where several conditions combined, and it is worth using in parallel rather than instead.
The test for whether you have reached a root cause is practical rather than philosophical: if you correct this condition, does the event stop being possible — not just here, but everywhere the same condition exists? If the answer is no, keep going. If the answer is that you would need to correct something outside your control, you have gone one step too far and the previous level is where the action belongs.
A finding without an owner and a date is a suggestion
The most common failure after a good analysis is administrative. The investigation is performed properly, a document is produced, the document is filed, and eighteen months later a similar event occurs in a similar way. Nobody was negligent. The analysis and the corrective work simply lived in different places, so nobody could tell whether the finding ever became anything.
A corrective action needs four properties to survive: a named owner rather than a department, a due date, a link to the asset or process it concerns, and a closure step that requires evidence rather than a checkbox. Miss any one and the action becomes a line in a spreadsheet that gets marked complete during the next audit preparation.
The link to the asset matters more than it first appears. When every investigation joins the permanent history of the equipment involved, the next person who opens that record sees what has already happened to this machine — and a pattern of three similar events across two years becomes visible in a way no individual report ever makes it. That is also what turns a stack of investigations into a trend rather than an archive.
Score the risk. Then score it again.
Scoring failure modes after a root cause analysis is common practice. Severity, occurrence and detection are rated, multiplied into a risk priority number, and used to decide what gets attention first. Most organisations that run structured investigations do this.
Almost none re-score afterwards. And the second score is where the value is, because without it you know what you intended to do and not how much doing it actually helped. A corrective action that reduces detection difficulty from eight to three has changed something real. One that leaves all three numbers where they were has produced activity rather than improvement, and the risk register cannot tell you which of those you have.
| Question | Answered by |
|---|---|
| What should we fix first? | The initial score, ranked. |
| Did the fix work? | The re-score after the action closed. Nothing else answers this. |
| Is the programme improving? | Total risk reduced across all closed actions over a period. |
| Which actions were theatre? | Closed actions with no change between the two scores. |
That last row is uncomfortable and worth having anyway. Every organisation closes some corrective actions that changed nothing, and the only way to stop repeating them is to be able to see them. Revision history on the scores makes the trend auditable rather than reconstructed, which is also what an inspector is looking for when they ask what effect the corrective action had — the point covered in the EHS audit checklist.
Do not wait for an injury
Near misses with high potential
The same conditions as an injury, minus the injury. Investigating these is the cheapest analysis available, because nobody was harmed and nobody is defensive.
Repeat conditions
The same hazard reported three times means the previous corrective action did not work. That is a finding in itself, and it points at the analysis rather than at the site.
Unplanned equipment failure
Reliability and safety share most of their root causes. A breakdown investigated properly frequently reveals a condition that would eventually have hurt somebody.
Successful recoveries
When something went wrong and the controls worked, understanding why is as valuable as understanding failure — and it is the only investigation nobody dreads being part of.
Investigation is the learn stage of the safety lifecycle, and it depends entirely on the detect stage feeding it. An organisation with excellent root cause analysis and poor near miss reporting is investigating only the events that hurt somebody, which is the most expensive possible source of learning.
How SmartX HUB handles investigation
Five whys and a cause diagram sit on the investigation record itself rather than in a free-text field, so the chain can be queried later and turned into a trend. Failure modes carry severity, occurrence and detection scores, and corrective actions carry their own re-scored values — so the system records not only what was done but how much it reduced the risk, with revision history making the trend auditable. Every investigation joins the permanent history of the asset involved, and findings become tracked work orders with an owner, a date and a closure step that requires evidence.
The seven stages of the safety lifecycle and where each one breaks.
From report to corrective action to closure, with the evidence attached.
Catching the condition before it becomes the event you investigate.
Find out whether your corrective actions worked
See how SmartX HUB scores risk before and after each corrective action, so the register measures improvement rather than intention.
Explore our blog for insightful articles, personal reflections and ideas that inspire action on the topics you care about.
Explore more information about cases and news.
Stay informed with the latest updates and in-depth insights.
RFID Tool Tracking & MRO — Common Questions
Answers to the questions we hear most from teams evaluating RFID tool tracking for maintenance and MRO operations. Have another question? Reach out through our support center.
What are the main challenges of MRO tool control?
How does RFID compare to barcodes, NFC, BLE or GPS for tools?
How do you set up an RFID tool tracking system?
How does RFID integrate with our existing MRO and ERP systems?
What if some tools are too small or the wrong shape to tag?
How does RFID ensure calibration and maintenance compliance?
How does RFID support predictive and preventive maintenance?
What keeps tools accessible during downtime or power loss?
Still have questions about bringing RFID tool tracking to your operation?
Talk to Our Team
