Why redundancy sometimes makes things worse
Backup systems that add failure modes instead of removing them.
Adding a second, independent system to stand in if the first one fails only improves reliability if the two systems really do fail independently of each other, and a great deal of redundancy that looks reassuring on paper quietly fails this test, either because the backup shares a hidden cause of failure with the system it is meant to protect against, or because the extra hardware needed to switch between the two has become a new point of failure that did not exist before anyone added a backup at all.
What is really going on
The appeal of redundancy rests on a genuinely sound piece of arithmetic: if one system has a small independent chance of failing and a second, entirely separate system covers the same job with its own equally small and genuinely independent chance of failing, the chance that both fail at once is smaller again, often dramatically so, than the chance of either one failing alone. This is why aircraft carry more than one hydraulic system, why data centres run more than one power feed, and why a critical bolt is sometimes backed up by a second, independently loaded path. The arithmetic depends entirely on that word independent, though, and independence is a much harder thing to actually build than it is to assume, since a backup system installed alongside a primary one very often ends up sharing something with it, the same power supply, the same enclosure, the same technician who services both on the same afternoon, the same design flaw copied into both because both were specified from the identical drawing. Once that shared thing exists, the two systems are no longer failing independently, they are two expressions of the same underlying risk, and the comforting arithmetic that justified adding the backup in the first place no longer describes what is actually happening. Nothing about the redundancy itself has to be built badly for this to happen, since both systems can individually meet every requirement placed on them and still share the one hidden dependency nobody thought to trace, which is exactly what makes a false redundancy so hard to catch during an ordinary design review that checks each system's own specification rather than asking what the two actually have in common.
The spare-tyre comparison
A spare tyre sitting in the boot of a car is redundancy in its simplest, most familiar form, a second wheel held in reserve specifically for the moment the primary one fails. What a great many drivers discover only at the worst possible time is that the spare has been slowly losing air for months or years, unnoticed precisely because it is the one tyre nobody ever has a reason to check, since it never touches the road and never gives any warning the way a slowly deflating road tyre eventually does through handling or a dashboard light. The spare was never actually as reliable as the tyre it was meant to replace, it only looked that reliable because it was never tested, and the very quality that made it a backup, sitting untouched and unused, is the same quality that let its own slow failure go completely undetected until the one moment it was finally needed and found wanting.
Why redundancy needs independence to actually work
A second failure mode this same example hides is what actually connects the spare to the car once it is needed, the jack, the wheel nuts, the tools, all of which have to work correctly for the redundancy to pay off at all. None of these exist in the primary system, they exist only because a backup was added, and each one is a genuinely new way for the overall job, getting the car moving again, to fail that simply was not present before anybody thought to carry a spare. This is a pattern that shows up well beyond spare tyres: a backup generator installed alongside a building's main power supply does not fail nearly as often from the generator itself as it does from the automatic transfer switch meant to detect the outage and hand the load across, a piece of equipment that exists purely because redundancy was added and that, being its own single point of failure, can quietly undo a large share of the reliability the second generator was bought to provide.
The number that matters here
Two genuinely independent systems, each reliable on its own the great majority of the time, combine mathematically into a system that fails only in the rare case both fail together, a combined reliability dramatically better than either system alone offers. The moment those two systems share even a modest chance of failing from the same underlying cause, that dramatic improvement collapses toward the reliability of the weaker single system by itself, since a shared cause can take out both at once regardless of how independently each one was designed to behave the rest of the time.
What this changes in practice
Reviewing a redundant design honestly means asking a harder question than whether a backup exists, it means asking specifically what the two systems actually share, power, environment, maintenance schedule, design lineage, or the switching hardware standing between them, since any one of those shared elements can quietly convert what looks like two independent chances of survival into one chance wearing a disguise. A backup that has never once been tested under the conditions it is meant to be relied on in has not actually been proven redundant at all, it has simply been assumed redundant, and the difference between those two states tends to stay invisible for exactly as long as the backup is never actually called upon, which is usually the entire span of time until the one moment it matters most. Scheduling a deliberate, periodic test of the backup on its own, disconnected from the primary system it normally rides alongside, is the one habit that actually closes this gap, converting an untested assumption into a genuinely checked fact, and an organisation that treats its backups as worth testing on their own schedule rather than only inheriting whatever attention the primary system already receives is, in practice, the organisation least likely to discover its redundancy was hollow at the exact moment it needed it to be real.