Replacing an electronics stack instead of patching it
Deciding to rebuild rather than continue correcting.
Replacing an electronics stack instead of patching it becomes the right call once the accumulated cost of every future patch, added to the cost of every patch already made, has grown larger than the cost of building the stack again from a clean design, and the hard part is almost never doing the arithmetic, it is noticing the point at which the arithmetic has already flipped.
Why the tenth patch can cost more than starting over
An electronics stack, the collection of boards, connectors, and wiring that make up a machine's control system, rarely fails all at once. It accumulates individual faults over time, a connector here, a marginal power supply there, a sensor interface that never quite behaved the way the datasheet promised, and each fault gets fixed as it appears with whatever targeted correction actually solves that one problem: a fly lead soldered in to bypass a bad trace, a filter capacitor added to quiet an interference issue that only ever showed up on one particular machine, a firmware workaround written to compensate for a timing quirk nobody had time to trace back to its actual electrical cause. Each of those fixes, judged on its own, is a perfectly reasonable use of an afternoon compared with redesigning the whole board around the fault. The trouble is that the stack these fixes are applied to gets more fragile and harder to reason about with every one of them, since a board carrying a dozen individual patches has a dozen more places a future change can interact badly with something that was never part of the original design, and a fault appearing on a heavily patched board takes measurably longer to diagnose than the identical fault would take on a clean one, purely because there is more accumulated history to rule out first.
The rebuilt-bicycle comparison
A bicycle ridden hard for years rarely needs replacing all at once either; it accumulates individual repairs instead, a new chain when the old one stretches, new brake pads, a replaced cable here and a trued wheel there, each repair entirely reasonable and each one cheap compared with buying a whole new bicycle on the spot. Carry on long enough, though, and a rider can look back at the accumulated cost of every individual repair and realise it has quietly added up to considerably more than a new bicycle would have cost outright, while the frame underneath all those repairs, the one part nobody thought to replace because replacing a frame basically means buying a new bicycle anyway, has aged the entire time without anyone directly addressing it. No single repair along the way was ever the wrong decision in isolation, and that is exactly what makes the total so easy to lose track of, since nobody ever sees the running total, only the next individual repair, which always looks small and always looks reasonable next to the alternative of buying an entirely new machine.
Why each individual fix looks reasonable and the total does not
The reason this pattern is so easy to fall into is that every decision along the way is being made under the same comparison, this one patch against a full replacement, and a single patch is almost always cheaper than a full replacement when it is judged against that specific replacement in isolation. What almost never gets compared directly is the running total of every patch made so far, plus every patch the stack is still likely to need, against that same full replacement, and that second comparison is the one that actually matters, because it is the one a full replacement would genuinely be judged against if anyone stopped to add it up. A stack with a long history of individual fixes is also carrying a second, less visible cost alongside the direct expense of each patch, the compounding difficulty of diagnosing anything new on a board that no longer quite matches its own documentation, since a fly lead or a workaround added under time pressure often never makes it back into the official schematic at all. This connects directly to the version control and traceability discussed elsewhere in this set: a patched board that was never formally revised is a board whose real, physical configuration has quietly drifted away from the one record everybody still trusts, and every hour spent later reconciling the two is itself a cost the original quick fix never appeared to carry.
The number that matters here
A stack that has already absorbed several rounds of individual patches can end up costing more in accumulated rework, each fix consuming both an engineer's time and a portion of the board's own remaining physical space for yet another bodged connection, than a full redesign would have cost from the outset, without the total ever being visible at any single point along the way it was actually being spent.
Why this matters in practice
The practical fix is not to refuse every patch on principle, since plenty of individual faults genuinely are cheaper to fix directly than to justify a full redesign over, it is to keep a running account of the patches made to a given stack rather than judging each one purely in isolation, and to revisit the comparison against a full rebuild deliberately, on a schedule, rather than only when a patch finally becomes impossible to make at all. A stack that has needed three or four significant patches inside a single year is very often already past the point where the next patch is genuinely the cheaper option, even though, examined entirely on its own, that next patch will still look exactly as reasonable as every one before it did. Building that review into a fixed schedule rather than leaving it to judgment in the moment matters precisely because judgment in the moment is the thing already shown to fail at this comparison, since it is always being exercised on the smallest, most recent piece of a much larger picture, and never on the picture as a whole.
My key error with this
The stack had grown the way these things do, with a general purpose computer doing the heavy work, a smaller board added later to handle the things the first one was awkward at, and a collection of adapters and cabling between them that each solved a real problem at the moment it appeared. Every patch was defensible on its own and the assembly as a whole had become the main source of faults, most of them at connectors, most of them intermittent, and all of them expensive to chase in the field. Replacing it rather than continuing to correct it was the right call for reasons well beyond tidiness, because consolidating onto a single compute module with microcontrollers out at each actuator meant far less cabling to shake loose, and moving to parts that did not need active cooling meant the enclosure could finally be sealed properly instead of being ventilated to keep fans fed. What replaced the belief is that a patch is a loan rather than a fix, and that the moment worth rebuilding is when the connections between the parts have become less reliable than any of the parts themselves.