← Back to Archive

Why accidents are usually organisational rather than technical

Why the technical cause is rarely the whole explanation.

A major accident almost always has a clean, identifiable technical cause once investigators finish taking it apart, and yet the technical cause on its own rarely explains why the failure actually happened, because the part or the material or the seal that gave way was usually a known risk that had already been noticed, written down and argued about by somebody long before it actually failed.

Challenger: The Final Flight spends four full episodes establishing exactly this, that the O-ring problem was not a hidden defect nobody had thought to check, it was a documented concern raised again by engineers the night before launch, and watching the argument play out on screen is uncomfortable precisely because the technical explanation was never really in question, the organisational one was.

The mechanism behind an organisational failure

A technical root cause answers the narrow question of what physically broke, and that question is usually the easiest one for an investigation to settle, since a failed part leaves physical evidence a laboratory can examine directly. The harder and more revealing question is why a risk that somebody inside the organisation had already identified was still running at the moment it turned into a failure, and answering that question means looking past the part itself and into how information about that risk moved, or failed to move, through the layers of people who could have acted on it. A warning raised by an engineer close to the hardware has to travel upward through people further from the hardware and closer to schedule, budget and reputation, and at every one of those handoffs the warning can be softened, reframed as routine, or simply not weighted as heavily as it was by the person who first raised it, until a genuinely serious technical concern arrives at the point of decision looking like one item on a long list of ordinary open issues rather than the one that matters most. None of this generally happens through any single dishonest act along the way, which is part of what makes it so hard to prevent, since each individual person softening the warning by a small amount is usually acting in good faith on the information immediately in front of them, and it is only the accumulated effect of several honest, small reframings in a row that turns an urgent concern into a routine one by the time it reaches whoever actually has the authority to stop the schedule.

The smoke-alarm comparison

A smoke detector that goes off from ordinary kitchen smoke often enough eventually gets treated differently by the household it protects, not because anybody consciously decided the alarm no longer mattered, but because every false alarm that turns out to be nothing but burnt toast quietly recalibrates how urgently the next alarm gets taken. Reaching up to silence it becomes a habit performed on the way back to whatever was happening before, rather than a genuine check of what triggered it, and by the time an alarm finally does go off for a real fire, it receives the same automatic, half-attentive response as every false one before it, the household's collective sense of the alarm's meaning having drifted a great deal further from "danger" than anybody in the house would have admitted if asked directly. A technical warning inside a large organisation can drift the same way, each previous instance where the same warning was raised and nothing bad happened afterward quietly teaching everyone involved that this particular warning tends not to matter as much as it sounds, right up until the one time it does.

Why a known risk stops being treated as a risk

This drift has a name in the safety literature, normalisation of deviance, and it describes something more specific than simple carelessness: a genuine, documented departure from the original safety margin that is observed, survived, and then folded into what counts as normal rather than treated as evidence the margin itself needs revisiting. Each time a part performs its job while showing some sign of the known problem, without that problem actually causing a failure, the organisation collectively updates its own sense of how dangerous the problem really is, and it updates it downward, since surviving the exposure feels like evidence the margin was adequate after all rather than evidence the margin was simply not yet exhausted. This is precisely the opposite of how the same evidence would be read by someone encountering it for the first time, uncoloured by the history of previous flights or previous cycles that also showed the problem and also survived, which is exactly why an outside review, brought in specifically because it has no such history to be reassured by, so often catches a risk the people closest to it have gradually stopped seeing.

The number that matters here

An organisation that has survived a known, documented departure from its own safety margin several times running without incident has not actually gathered several data points confirming the margin holds, it has gathered several data points showing only that the margin has not yet been exhausted, a distinction easy to state and remarkably hard to act on once each individual survival has already been logged internally as a quiet success.

What this changes in practice

Investigating an accident thoroughly means asking not only what broke, but who knew, how far up the organisation that knowledge travelled before the failure, and at what point along that path the original engineer's sense of urgency stopped being carried forward with it, since the technical fix that follows a well-run investigation is usually the easy part, and the organisational fix, changing how a raised concern actually moves and how much weight it keeps as it moves, is the part far more likely to determine whether the same failure happens again somewhere else in the same organisation, wearing a different part's name. A report that stops at the failed component and recommends only a stronger seal, a thicker wall or a tighter tolerance has answered the question an outside observer asks first and skipped the one that actually explains why the organisation was running that risk in the first place, which is exactly the gap the next two articles in this set go on to explore from two different directions.

More on Failure as a system