The simulation that agreed with the test for the wrong reason
Two errors cancelling, and how long it took to notice.
A simulation agreeing with a physical test for the wrong reason happens when two separate mistakes in the model happen to cancel each other out at the specific condition being compared, producing a match that looks like validation while proving nothing at all about whether either underlying mistake has actually been found.
Two mistakes pulling in opposite directions
Validation normally works by comparing a model's prediction against a physical measurement and treating agreement as evidence the model is sound. That logic assumes the model has only one source of error at a time, or that its errors all push the answer the same way, and neither is guaranteed. A model can carry two separate mistakes, a boundary condition set slightly too high and a material property set slightly too low, say, that shift the final answer in opposite directions by roughly matching amounts. At the exact condition being compared against a test, those two errors can cancel almost completely, and the comparison shows excellent agreement while nothing about the model has been confirmed.
The numbers really do agree, often closely enough to look like the kind of validation success a project celebrates. The comparison cannot see that the agreement comes from two wrongs landing in the right place. Change the tested condition slightly and the two errors, no longer scaled to cancel so neatly, pull apart again, revealing a mismatch hidden by a lucky alignment at the one point anybody checked.
This failure is hard to catch because it produces exactly the evidence everyone is hoping to find. A validation exercise is usually judged a success the moment simulation and test line up, and there is rarely any pressure to keep digging after that, since digging looks like second-guessing good news. The cancelling errors hide behind the very success criterion that was supposed to catch them.
The more parameters a model carries, the more places an independent error can hide and the more combinations of errors might offset one another at a single condition. A simple model built from a few inputs has few ways to fail by lucky coincidence, while a detailed model carrying dozens of boundary conditions, material properties and modelling choices has a great many, so sophistication alone earns a model no extra trust.
Two wrong clocks showing the right time
Take two clocks, one set ten minutes slow but running five minutes an hour fast, the other set ten minutes fast but running five minutes an hour slow. Two hours later both show exactly the correct time, and anyone glancing at them at that moment, seeing them agree with each other and with the truth, would feel reassured. Neither clock has been fixed. An hour after that, one is five minutes fast and the other five minutes slow, ten minutes apart, and the mismatch that was present all along is visible again. The clocks were only passing through the same point at the same instant on their way to being wrong in opposite directions.
A model tested at one condition is those two clocks read at the one moment they coincide. The boundary condition that was set too high and the material property that was set too low each have their own drift, and the test point happened to be the instant at which those drifts summed to zero. Move the test to a higher load, a different speed or a warmer day, and the two errors scale by different amounts, so the model that matched perfectly a moment ago now misses by a margin nobody saw coming.
Bringing a third clock into the comparison makes the coincidence far easier to catch, since three independently drifting clocks landing on the same reading at once is much less likely than two managing it. A model checked against several independent measurements, taken under different conditions, is harder to fool by cancellation for the same reason, since each new measurement is another clock that has to happen to agree.
Agreement across a spread of conditions
A validation run at a single test condition offers one data point, and one data point cannot distinguish an accurate model from two errors that cancel at exactly that condition. Running the comparison across a spread of conditions tests whether the agreement is real, because errors of this kind are rarely arranged by nature to cancel across a whole range. Repeated agreement as conditions change would need the errors to keep cancelling precisely at every step, which becomes less plausible the wider the range gets.
That wider spread costs time and equipment access, which keeps the shortcut of trusting one comparison tempting, and it is also the fastest way to mistake a lucky cancellation for a validated model. A single clean comparison is treated as encouraging, and validation against a varied set of conditions, beyond the one that happened to be convenient to test first, is the standard the discipline holds itself to.