Running a machine at full capacity slows everything down
The counterintuitive result that governs every production system.
Running a machine at full capacity slows everything down because a resource loaded right up to its limit has no slack left to absorb the small, ordinary variation in when work arrives, so the queue in front of it grows sharply, and ever more sharply, the closer its workload creeps toward one hundred per cent.
A packed motorway moves fewer cars
A motorway carries more cars per hour at a brisk, comfortably spaced density than when every lane is packed bumper to bumper at what looks, to a driver stuck in it, like the road's absolute capacity. Once vehicles are spaced so tightly that every driver is constantly adjusting to the car ahead, a single touch of the brakes ripples backward as an amplifying wave, each following driver reacting a moment later and braking a little harder, and that wave can bring a stretch of road to a stop long after the original slowdown has cleared.
The gaps between vehicles on the brisk road are the slack that lets drivers absorb small variations in speed and spacing without setting off that wave. Watched from above, the moment density crosses from busy-but-flowing into fully packed, the count of cars passing a fixed point each hour falls, even though the road holds more vehicles at any instant than it did minutes earlier.
Why the queue curve turns upward
A resource running at seventy or eighty per cent of what it could process has room to absorb a cluster of arrivals, because the gaps between earlier arrivals left it with unused time to spend catching up. A resource running at ninety-nine per cent has almost none of that room, so any cluster landing on it joins a queue that keeps growing for as long as the cluster keeps arriving faster than it can be cleared.
The queue therefore stays fairly tame across most of the range and then climbs steeply near the ceiling. For a single machine with randomly timed arrivals and job lengths (the standard textbook case), the average wait in line equals the job time multiplied by the busy fraction divided by the idle fraction. At eighty per cent busy, a job waits four job-lengths before starting. At ninety per cent it waits nine, and at ninety-nine per cent it waits ninety-nine. The step from ninety to ninety-nine per cent adds a tenth to the output and multiplies the waiting by eleven.
The reason lies in the idle fraction at the bottom of that ratio. A queue can only shrink during moments when the machine would otherwise have been free, so the idle time is the whole budget for recovering from a bad cluster. At eighty per cent busy there are twelve idle minutes in every hour to drain a backlog; at ninety-nine per cent there is barely half a minute. Real factories rarely match the textbook assumptions exactly, but the shape of the curve, flat for most of the range and steep near the top, is the same, just as the motorway's traffic holds up well until the last gaps close.
A manager watching a utilisation report climb from seventy per cent toward the high nineties, expecting the queue to climb at a similar steady rate, is working from the wrong picture. The report shows a modest, reassuring rise in busyness at exactly the point where the queue turns sharply upward.
Protecting the constraint with deliberate slack
Running a critical resource somewhat below its theoretical maximum, instead of scheduling it to be busy every available minute, keeps the margin that stops queues from growing sharply the moment ordinary variation shows up. A plant that keeps its most heavily loaded machine slightly under its ceiling on purpose often moves more work through the whole system per day than one that schedules that machine at its limit and then spends the week fighting the queues that decision produces.
The theory of constraints treats this directly, arguing that a production system's total output is set by its single tightest resource, and that protecting that resource's steady flow is what raises how much the whole line produces. The instinct to treat every idle minute as pure loss runs against this result, since the idle minutes absorb variability that would otherwise surface downstream as a queue nobody planned for.
The margin belongs specifically at the constraint. A queue there delays everything downstream, while the same slack given to a machine with plenty of spare capacity is unused time that nothing was waiting on. Working out which resource deserves the margin is a real piece of judgement, and getting it wrong in either direction, protecting a machine that was never the constraint or squeezing the one that was, produces the same result: a plant that looks busier everywhere and ships no faster than before.