The failure that only ever happened in the field
A fault that could not be reproduced anywhere it could be observed.
A fault that shows up reliably in real field use and then refuses to appear on a bench, no matter how carefully the bench test is built, almost always turns out to depend on several specific conditions arriving together at once, a combination the field naturally produces every so often and a simplified bench test, built deliberately to isolate one variable at a time, essentially never stumbles onto by chance.
Why a bench test cannot recreate what the field actually does
A bench test earns its usefulness precisely by simplifying the real world down to a small, controlled set of conditions that can be varied one at a time and measured cleanly, holding temperature steady while checking vibration, or holding load steady while checking temperature, because that discipline is what lets a specific cause be matched to a specific effect with any confidence at all. The field offers no such courtesy, delivering temperature, vibration, load, humidity and electrical noise all at once, in a shifting combination that changes from one hour to the next and is never held still for anybody's convenience. This is not a flaw in how bench tests are built, it is the entire reason they are useful at all, since a test that varied everything at once, the way the field does, would make it impossible to say afterward which of the many changing inputs had actually caused whatever result was observed, and an engineer would be left with a confirmed failure and no way to trace it back to a specific cause worth fixing. A fault that genuinely requires several of those conditions to line up simultaneously, a particular temperature combined with a particular vibration frequency combined with a particular load, will simply never appear on a bench that is only ever varying one of those inputs while holding the rest at a convenient, steady baseline, because the bench was designed from the outset to prevent exactly the kind of simultaneous, uncontrolled combination the fault actually depends on. A bench built to isolate variables is, in effect, betting that the failure mechanism responds to each input independently, an assumption that holds for a great many faults and is exactly why the bench remains the right first tool to reach for, but the bet fails outright the moment the mechanism only responds to several inputs acting together, since holding the others steady cannot substitute for actually letting them move at the same time.
The bicycle-creak comparison
A bicycle can develop a faint, specific creak that only appears when a rider stands on the pedals and pushes hard climbing a hill, the rider's full body weight and a sharp burst of torque combining in a particular way through the frame and drivetrain at that exact moment. A mechanic examining the same bike on a repair stand, spinning the pedals gently by hand with no rider's weight resting on them and no real torque passing through the chain, can check every part in turn, the bottom bracket, the pedals, the seat post, and hear nothing at all, not because the fault has vanished but because the repair stand was never subjecting the bike to the specific combined load that produces it in the first place. The creak was always real, and the bike was never actually fixed by the inspection that found nothing wrong, it was simply never tested under the one condition that had ever revealed the fault to begin with. A rider handing the bike over with nothing more specific to say than "it creaks sometimes" makes the mechanic's job considerably harder still, since a vague description gives no clue which of the many possible combined loads the creak actually depends on, and a mechanic working from that description alone has little choice but to guess at which component to suspect first, often the wrong one, before ever getting close to recreating the actual condition that produces the sound. Even a stationary trainer, which lets a mechanic stand on the pedals and push down with something close to a rider's real weight, still leaves out the small side-to-side sway of the frame that happens when a bike is free to lean under a climbing rider, so a fault depending on that sway as well as the load can survive a trainer test exactly as it survived the repair stand, passing every check that never happened to reproduce the one missing ingredient.
Why some faults need several conditions to align at once
The deeper reason this kind of fault resists reproduction is that it depends on an interaction between several inputs rather than on any single one crossing its own threshold alone, and an interaction of this kind is invisible to a test plan built around checking each input separately. A component can pass a pure vibration test, pass a pure thermal test, and pass a pure load test individually, and still fail in the field the moment all three happen to be present together at levels that would each have been harmless on their own, because the fault mechanism itself, whatever it turns out to be, may only activate once the combined effect of several inputs crosses a threshold that no single input reaches by itself. Chasing a fault of this kind by improving one bench test at a time, a stronger vibration table, a wider thermal chamber, rarely helps, since the missing ingredient was never the severity of any one test, it was the absence of the other conditions arriving alongside it. What actually closes the gap is a test built deliberately to run several of those conditions together rather than in careful isolation, accepting the loss of clean, single-cause attribution in exchange for a genuine chance of recreating the combination the field has already shown to be dangerous, a trade that only makes sense once single-variable testing has already been tried and has already failed to catch anything.
The number that matters here
A fault that depends on several specific conditions aligning at once can occur in genuine field use only once every several hundred hours of running, rare enough that a bench test lasting a handful of hours, however carefully instrumented, has essentially no realistic chance of stumbling onto the same combination purely by accident, which is exactly why a fault this rare in the field can still be, in effect, guaranteed never to appear on the bench at all. That gap between the field's natural rate and the bench's available running time explains why an engineer can watch a bench come back clear and still be wrong to treat the fault as fixed, since the bench was never actually capable of encountering the combination in the first place.
My key error with this
We had thought carefully about cooling the compute, because the processor is the part with a published thermal envelope and a fan already attached to it, and in the field that part behaved exactly as intended. What failed instead were the motor controllers and the small microcontrollers sitting alongside them, which I had never really considered as thermal components at all, since individually they dissipate very little and none of them had ever run warm on a bench. Enclosed in a sealed box, in the sun, with the motors working continuously against real ground rather than against a test load, they climbed steadily and began behaving erratically in ways that took a long time to attribute to temperature at all, because the symptom looked like a control fault rather than a heat fault. What replaced the belief is that thermal design is a property of the whole enclosure rather than of the components that advertise a heat problem, and that the parts most likely to catch you out are the quiet ones nobody assigned a cooling budget to.