Models are checked against reality, not the reverse
Validation, and the direction the evidence has to flow.
Models are checked against reality, not the reverse, because a simulation can only ever be internally consistent with itself, and the one thing internal consistency cannot supply is proof that the model actually describes the physical part it claims to represent.
Vincenti's What Engineers Know and How They Know It traces how aeronautical engineers spent decades arguing over what should count as proof that a piece of theory could be trusted, and the argument, more often than not, came back to the same question: had it been checked against something real.
A closed system of assumptions
A simulation, however carefully built, is a closed system of assumptions: a geometry, a mesh, a set of equations and a set of boundary conditions, all chosen by a person and none of them verified by anything inside the software. The solver can confirm that the numbers it produced satisfy the equations it was given, and it has no mechanism at all for confirming that those equations, that geometry or those boundary conditions correspond to the real part sitting on a bench. Consistency is the only one of the two properties a model can demonstrate about itself.
Validation supplies the missing half by comparing a model's prediction against an independent measurement of the real thing, a physical test, an instrumented prototype, a known result from a similar part built before. Only that comparison brings in evidence from outside the closed system.
An oven dial that is wrong the same way every time
An oven's thermostat dial can be perfectly consistent with itself. Set it to the same mark twice and the oven reliably reaches the same internal state both times. That repeatability says nothing about whether the dial tells the truth about the temperature inside, because a dial can be consistently wrong as easily as consistently right, always landing on the same incorrect number. The only way to know is to put an independent thermometer inside and compare its reading against the dial.
A simulation checked only against its own convergence history is the dial checked only against itself. It can land on the same answer every time the same inputs are run and still describe a stress or a flow that an independent measurement of the real part would contradict, in the way an oven that reliably burns every dish at the same setting is perfectly consistent.
Calibrating with one set of data, checking with another
The thermometer has to be one the model was never allowed to see or tune against. Calibration adjusts a model's free parameters until it matches known data, and comparing the model against that same data afterwards proves only that the calibration worked, as circular as checking the dial against itself. Validation checks the calibrated model against a further, independent set of measurements taken to test it, and keeping those two activities separate is most of what makes the comparison worth anything.
Physical testing is also the slower and more expensive half of this exchange, needing a real part, real instrumentation and real time on a test bench, which is why it is the step most likely to be trimmed under schedule pressure, though it is the only one that can catch an error the model's internal consistency would never reveal. A model validated across a range of conditions earns trust within that range, and cautiously a little beyond it, while one never checked against anything outside itself earns it nowhere, however many iterations it has run.
Validation therefore belongs in the schedule as a required step. The decades of argument Vincenti documents among aeronautical engineers keep returning to the same conclusion, that a theory earns trust in the direction the evidence flows, from the physical test toward the model. A schedule that treats the physical test as a formality to confirm what the simulation already proved lets confidence build up around a model well ahead of the evidence that was supposed to earn it.