← Back to Archive

Why tolerance stacking is usually done wrongly

The two methods, their assumptions, and when each one applies.

Tolerance stacking is usually done wrongly not because engineers pick the wrong arithmetic, but because a single method, almost always whichever one the last project happened to use, gets applied as an unquestioned house habit to every new stack, whether or not that stack's production volume, process maturity and consequence of failure ever matched the assumptions the method was built on.

Two methods answering two questions

Worst-case and statistical stacking are each the right tool for a specific set of conditions, and using either outside those conditions produces a wrong answer even when the arithmetic is carried out perfectly. Worst-case stacking answers the question "will every single unit ever built definitely work." That guarantee matters enormously when failure is unacceptable, or when so few units will ever be made that there is no population for statistics to describe. Statistical stacking answers "will the great majority of a large, well-behaved production run work," and that question only has a trustworthy answer once real data or real confidence about the process's behaviour exists to support it. Mistaking one question for the other is the failure this article's title points at.

A wedding cake and a tray of biscuits

Baking a single elaborate wedding cake for one occasion leaves no room to rely on averages. Every layer has to come out right, because there is no batch of other cakes for a slightly sunken one to be balanced against, and that is the position worst-case stacking defends, a design with no population to lean on. Baking a hundred identical biscuits from one batter is a different situation. A few running slightly over or under size barely register, and the batch looks consistent because the ordinary, independent variation between biscuits mostly cancels across so many of them, precisely the assumption statistical stacking is built on.

Applying the wedding-cake mindset to the biscuits wastes effort demanding a perfection no single biscuit needed. Applying the biscuit mindset to the wedding cake gambles the one occasion that mattered on an averaging effect a single cake never had the numbers to benefit from.

Even the biscuits only get that cancellation if their variation is random. A domestic oven with one persistently hotter corner makes every biscuit placed there come out a little larger, batch after batch, a systematic bias repeating in one direction. A stack of parts machined on one worn tool, or measured against one subtly biased gauge, fails the statistical assumption in exactly the same way.

Short chains gain little

Statistical cancellation needs enough independent terms in the stack for aligned extremes to become unlikely, and the benefit grows with the count. For a chain of equal tolerances, the statistical result is the worst-case result divided by the square root of the number of terms. Two dimensions therefore bring the combined tolerance down only to about seven tenths of the worst-case figure, and three to under six tenths. Ten independent dimensions bring it to about a third, and twenty-five to a fifth.

A two-term stack gains a modest margin while carrying the real risk that this particular short chain lands unfavourably. A long chain of a dozen or more independently sourced dimensions has enough terms for the statistical argument to carry real weight. The same method, applied honestly, can therefore be right for one stack on a drawing and wrong for a shorter stack a few lines below it on the same page, purely because of how many independent terms each contains. Counting the terms before choosing the method takes a minute and settles much of the decision on its own.

How a house default survives

Almost nobody reasons deliberately from first principles to the wrong method. The failure is quieter: a template inherited from an earlier project, a calculation spreadsheet built once and reused for every drawing since, a habit nobody on the current team remembers choosing. The sheet typically reports a single combined figure with no note of which method produced it, so a reviewer reading the result has nothing prompting the question. The default also survives because it rarely causes an obvious disaster. An over-cautious worst-case habit merely costs more than it needed to and never produces a visible failure. An under-cautious statistical habit reveals itself only on the rare occasion a short or immature stack finally lands badly.

Both directions are easy to carry for years without correction, quietly costing money in one case or quietly accumulating risk in the other, until the cost or the failure grows large enough for someone to ask where the method came from.

Choosing dimension by dimension

The patterns above are guidelines, and particular cases overturn them. A short stack with severe consequences for failure, or a long stack built from a process too immature to trust statistically, can each go against the general rule. A single stack need not commit wholly to one method either. The few dimensions in a chain that must never fail can be held to worst-case limits while the less consequential dimensions around them get statistical treatment, a hybrid that asks the harder question dimension by dimension.

The discipline being argued for is asking, explicitly for every stack, whether the volume, the maturity of the process and the cost of being wrong match the assumptions the chosen method depends on. A drawing's tolerance stack deserves the same written justification as any other engineering decision on it, and a method chosen and recorded deliberately can be questioned and corrected later in a way an inherited habit never is.

More on Tolerance