Repeatability, and how much noise is normal
Separating a real effect from the scatter around it.
Repeatability describes how tightly a result would cluster if the exact same test were run again with nothing deliberately changed, and knowing how much that natural scatter usually amounts to is the only way to tell whether a difference between two results reflects something real or is simply ordinary noise.
Why an unchanged test still gives different answers
No physical measurement, however careful, comes back identical every time it is repeated. Small, uncontrolled variations creep in from everywhere: a slightly different ambient temperature, a fractionally different grip or angle, tiny differences in a material sample that nominally came from the same batch, rounding in whatever instrument is doing the reading. None of these amounts to much on its own, but added together they produce a real spread of results around the true underlying value, a spread that exists even when the thing being measured has not changed between one run and the next.
That spread has a rough size that belongs to the measurement itself rather than to anything being tested. Knowing it in advance, by repeating the same unchanged test before ever changing a variable on purpose, is what lets a later result be judged sensibly. A difference between two results only means something once it is large compared with the scatter the same setup would have produced anyway.
A free throw that lands left of the last one
A basketball player shooting the same free throw, with identical form, twenty times in a row does not land the ball in the same spot every time. Some shots drop cleanly through, some catch the rim, some miss to one side, purely from the ordinary variation in a human body repeating a practised action. One shot landing slightly left of the last is simply what a distribution of shots from an unchanged technique looks like, and only a shift across a whole batch of shots is worth reading as a real change in technique.
A coach who changed one thing about a player's form and judged it on whether the very next shot went in would be fooled constantly, rewarding a lucky bounce and punishing an unlucky one. Judging a change in a manufacturing process from a single pair of test results, one either side of the change, is exactly as unreliable.
Five repeats before touching anything
Taking the same unchanged setup through five or more repeats before altering anything gives an honest estimate of how much spread to expect from noise alone, and that estimate is what any later comparison should be measured against. A single before-and-after pair is one of the weakest forms of evidence available, since it cannot separate a real shift from an unlucky pair of ordinary measurements drawn from the same distribution.
Running a few repeats on both sides of the change, and comparing the spread within each group against the gap between the two groups, turns a plausible anecdote into something that can be trusted. An effect worth acting on should stand well clear of the repeat-to-repeat spread, because a difference that only just edges above the noise is exactly the kind of result most likely to vanish the next time the test is run. The extra runs are almost always cheaper than acting on noise mistaken for an effect.
Repeatability against reproducibility
A spread estimated from a handful of repeats is itself only an estimate, so a figure calculated from three or four runs deserves less confidence than one calculated from ten or twenty. There is also a meaningful difference between repeatability, how much a result varies when the same operator repeats the same test on the same setup, and reproducibility, how much it varies when a different operator, a different day, or a different piece of equipment is involved.
A change that looks real against a tight repeatability figure can shrink back into the noise once reproducibility is accounted for, because a second source of variation has now been let into the comparison. Treating the two as interchangeable is one of the more common ways a result that looked solid in one lab quietly fails to hold up in another.