← Back to Archive

Testing many things at once

Factorial designs, explained without notation.

Testing many things at once means deliberately running every combination, or a carefully chosen fraction of every combination, of several factors together in one organised set of experiments, in place of exploring each factor separately and hoping the results can simply be added up afterwards. This approach is usually called a factorial design, and despite the formal-sounding name the idea behind it is closer to common sense than the one-at-a-time habit it replaces.

Four loads of washing

Deciding how much detergent to use and what temperature to wash at could be tested by running a load with extra detergent at the usual temperature, then a separate load with the usual detergent at a higher temperature, and comparing each against a normal load. A factorial approach runs all four combinations across a single weekend: normal detergent cool, extra detergent cool, normal detergent hot, and extra detergent hot. Only once all four results sit side by side does it become possible to see whether extra detergent helps more at one temperature than the other, a pattern that comparing separate pairs of loads one change at a time would never reliably surface.

Running all four loads together also keeps the conditions steady, the same machine, the same week, the same batch of towels, where separate tests run days or weeks apart would each drift slightly. The comparison between combinations comes out cleaner as well as more complete.

Every run counts toward every answer

A factorial design starts from the same handful of factors a one-at-a-time study would use, but it visits every combination of them at whatever settings, usually just two, have been chosen for each. Two factors at two settings give four combinations, three give eight, and the count doubles with each additional factor, with every combination given an actual run.

Once the runs are complete, the effect of any one factor is found by comparing the average result with that factor at its higher setting against the average at its lower setting, taken across every combination of the other factors. Because the other factors were varied too, the comparison cannot be thrown off by an unrepresentative baseline the way a one-at-a-time study can. Because every combination exists in the data, any interaction between two factors appears directly as a quantity that can be calculated, where a one-at-a-time study would have left a gap in its coverage.

The four loads of washing show how this works. The detergent effect is the average cleanliness of the two extra-detergent loads minus the average of the two normal ones, one measured cool and one hot on each side. The temperature effect uses the same four loads, split the other way. The interaction is whether the detergent's improvement in the hot pair differs from its improvement in the cool pair, and if it does, that difference is itself the answer the weekend was run to find.

Every run also contributes to every calculated effect, which is why a factorial study extracts more information from a given number of runs than a sequential one. Four factors at two settings need sixteen runs, and those sixteen deliver the separate effect of each of the four factors plus all six two-factor interactions between them, ten results from one set of runs. A one-at-a-time study of the same four factors would spend five runs to get four numbers, with no interaction information in any of them.

Where interactions are plausible

Every conclusion from a factorial study about one factor has been checked across the full range of settings the other factors were run at, which gives the results a kind of confidence a sequential study cannot offer. That matters most wherever an interaction is even plausible: two chemicals whose combined effect might differ from the sum of their separate effects, two dimensions on a part whose tolerances might combine in how a joint seats, two settings on a machine whose joint effect on a finished part neither setting alone would predict. Running the full set of combinations costs more upfront than a one-at-a-time study, and in exchange it leaves nothing important unexplored merely because it sat off the single narrow path a sequential study was willing to walk.

A thousand runs for ten factors

The doubling that makes the design complete is also its biggest limitation. Ten factors at two settings each need 1,024 runs to cover every combination, well beyond most budgets and schedules. In practice full factorial studies rarely go beyond five or six factors, which already means 32 or 64 runs, and even those are only affordable when a single run is quick and cheap, a short simulation or a small physical test, and out of reach when each run is lengthy or destroys the part.

That limit is the reason a second family of methods exists, fractional designs that deliberately give up some of the full design's completeness, usually the ability to separate certain interactions cleanly from each other, in exchange for a fraction of the runs. A half-fraction of the six-factor study, for example, needs 32 runs in place of 64. Choosing between a full design and a reduced one is a real trade, and what is being given up for fewer runs needs to be understood and accepted at the outset, before the results arrive and it turns out that the one interaction that mattered was among those deliberately blurred together.

More on Experiments