Building an instrument to find a fault instead of fixing the fault
Choosing to build a test platform rather than continue patching.
Once a fault has already resisted repeated patching without ever revealing its actual cause, the more productive move stops being another attempt at a fix and becomes a deliberate investment in a tool built for one purpose only, catching the fault in the act with enough data attached to it that its real cause can finally be seen.
Machines kept failing the same way in the field and nobody could reproduce it on a bench, and after enough rounds of patching whatever seemed most likely to be responsible, the honest answer was that nobody was actually fixing anything, they were guessing, more carefully each time but guessing all the same.
Why patching without a diagnosis eventually stops working
Patching a fault means applying a plausible fix based on a theory of what caused it, and the first few attempts at this are usually reasonable, since a plausible theory tested quickly is often cheaper than building anything more elaborate to confirm it first. The trouble starts once several plausible theories in a row have each been tried and each has failed to actually stop the fault from recurring, because at that point the pattern is no longer that one specific theory was wrong, it is that the whole approach of guessing a cause and testing the guess directly in the field has run out of good guesses to try, and continuing to patch past that point spends real effort without any longer being informed by anything beyond hope. Each individual patch also carries its own hidden cost beyond the labour of applying it, since a fix based on the wrong theory can mask the fault's symptom for a while without touching its actual cause, buying a stretch of apparent calm that ends the moment the real cause finds another way to express itself, often in a form that no longer resembles the original fault closely enough to be recognised as the same problem at all. Worse still, a patch that happens to suppress the symptom for a while without touching the cause actively destroys the very evidence that might otherwise have led somewhere useful, since a fault that briefly stops recurring looks, to anyone reviewing the record afterward, exactly like a fault that has actually been fixed, and the next round of failures then has to start the search over again with no memory of which theories the previous round had already quietly ruled out.
The garden-camera comparison
Something has been getting into a vegetable garden overnight, and the first response is patching the fence wherever the damage seems to have happened, filling one gap, reinforcing one weak post, only for the damage to keep appearing somewhere slightly different a few nights later. Filling gap after gap without ever actually seeing what is getting in treats every new hole as its own isolated problem, when the real cause, whatever animal is actually responsible, has never once been identified, only its most recent point of entry. Setting up a simple trail camera pointed at the fence line instead is a different kind of move entirely, spending effort not on the next patch but on finally seeing, directly and unambiguously, exactly what is getting through and how, and once that picture exists, fixing the actual problem, rather than its latest visible symptom, finally becomes possible in a way no amount of further patching ever could have delivered on its own. Nobody sets up the camera expecting an immediate answer either, since a fault or an intruder rare enough to have resisted every earlier fix is also rare enough that the camera may need to sit recording for several nights running before it finally captures the one visit that matters, a patient, deliberately unglamorous kind of effort that looks, from the outside, indistinguishable from doing nothing at all right up until the moment it actually pays off. Moving the camera every time a fresh gap appears, chasing the most recent damage the way the fence patching did, would waste the one advantage the camera actually has over patching, so the more disciplined approach is to leave it recording on whichever stretch of fence has been hit more than once, trusting that the animal responsible for a repeated pattern of damage is more likely to return to a place it has already found productive than to strike an entirely fresh spot on any given night.
Why an instrument is a different kind of investment than a patch
A dedicated instrument built specifically to catch a fault, whether that means a data logger carried into the field alongside the equipment it is monitoring, or a bench rig deliberately built to recreate several field conditions at once rather than testing them one at a time, costs real time and money before it produces anything useful at all, and that upfront cost is precisely why it gets deferred in favour of one more patch attempt for as long as any plausible patch still seems worth trying. What changes the calculation is recognising that a patch attempt with no diagnosis behind it is not actually cheap, it only looks cheap because its cost is spread thinly across many small individual attempts rather than concentrated into one visible bill, and once enough of those individual attempts have accumulated, their total cost can already exceed what the instrument would have cost from the very start, with the added disadvantage that none of that accumulated patching effort produced an actual diagnosis the way the instrument eventually will. Recognising the exact moment that calculation has already flipped, before the accumulated patching cost is obvious to everyone in hindsight, is the harder and more valuable judgement call, since committing to build an instrument too early wastes money chasing a fault a slightly better guess might genuinely have fixed outright, while committing too late has already spent more than the instrument would have cost several times over. There is no clean formula that names the crossover point in advance, since the cost of the next patch attempt and its odds of working are both genuinely uncertain going in, which is why the decision tends to fall to judgement rather than to a calculation anyone could simply look up.
The number that matters here
A team that has patched the same recurring field fault several times over without ever identifying its actual root cause has very often already spent more in accumulated patching effort than a dedicated instrument built specifically to catch the fault in the act would have cost from the outset, the difference being that the instrument's cost is visible and committed to all at once while the patching cost quietly accumulates in smaller pieces nobody ever adds up until well after the fact.
My key error with this
We had a run of sheet metal parts come back wrong, and none of the causes were interesting, since they were small omissions in the files we had sent, a duplicated line here, an open contour there, a radius that was smaller than anything the shop could actually produce. My response for a while was to look harder before sending, which works exactly as well as looking harder ever does, meaning it works until the day it does not. Writing a small tool to read the files and check them against the things that had bitten us was not a difficult piece of software, and it was never clever enough to correct anything, but it flagged the specific problems we had actually met before, it almost never raised a false alarm, and once it passed a file that file went through the shop cleanly every single time. What replaced the belief is that a recurring error made by a careful person is a process problem rather than an attention problem, and that the correct response is to build the check rather than to promise to concentrate.