PID explained without a single equation
The three corrections, described by what each one is trying to fix.
PID control is just three separate corrections added together, one reacting to how large the current error is, one reacting to how long that error has been sitting there uncorrected, and one reacting to how fast the error is changing, and every one of the three exists purely to fix a specific way that a simpler, single correction falls short on its own, without a single equation required to see why each one earns its place.
The mechanism behind proportional-integral-derivative control
The proportional term is the most intuitive of the three, correcting in direct proportion to how large the current error is, a bigger gap between where the system is and where it should be producing a bigger correction, a smaller gap producing a gentler one. On its own, proportional correction has an awkward weakness, since a correction that shrinks as the error shrinks tends to settle into a small but permanent gap rather than ever fully closing it, because by the time the error is tiny the correction being applied has become too small to finish the job against whatever steady resistance the system is fighting. The integral term exists specifically to close that permanent gap, by accumulating the error over time and adding an ever-growing correction for as long as any error remains, so that even a small persistent gap eventually gets pushed out entirely. The derivative term exists for a different reason again, reacting not to the size of the error or its history but to how quickly it is changing right now, applying a restraining correction when the error is closing in fast in order to keep the system from sailing straight past the target it was aiming for.
The hill-driving comparison
Cruise control on a hilly road is a clean, everyday example of all three terms working together. Climbing a hill, the car's speed starts to sag below the set target, and the proportional term responds immediately by pressing the accelerator harder in rough proportion to how far speed has fallen. If the hill is long and steady, proportional correction alone tends to leave the car cruising just slightly below the target speed indefinitely, since the small remaining error only produces a small remaining push on the accelerator, not quite enough to fully close the gap against the constant resistance of the slope. The integral term is what finally closes that lingering gap, quietly adding a little more accelerator for every additional second the car spends below target speed, until the accumulated correction is finally enough to bring the car exactly back to the set speed even while still climbing. Cresting the hill and starting downhill, speed now rises quickly, and the derivative term is what eases off the accelerator early, reacting to how fast speed is climbing rather than waiting for it to actually overshoot the target before doing anything about it, easing the car back toward the set speed smoothly instead of letting it swoop past and have to be reeled back in afterward.
Why all three are rarely needed at full strength
Real control problems vary enormously in which of the three terms actually matters, and a well-tuned controller is rarely running all three at equal strength. A system with very little steady resistance to overcome, a light mechanism with almost no friction to fight, may need almost no integral correction at all, since proportional control alone already closes the error most of the way on its own. A system where overshoot barely matters, or where the measurement itself is too noisy to trust, may be tuned with very little derivative correction, a problem explored in more depth later in this set. Understanding what each term is actually trying to fix, rather than treating all three as a fixed package that must always be present together, is most of what separates a controller that has been properly tuned from one that has simply had all three dials turned up out of habit.
One figure worth keeping in mind
A proportional-only controller left to run a system against a persistent load can settle into a steady-state error that never closes on its own, sometimes remaining a noticeable fraction of the original error indefinitely, while adding even a modest integral term is often enough to drive that same lingering gap to essentially nothing given enough time, which is why so many practical controllers include at least some integral correction even when derivative correction is judged unnecessary or left out entirely for the reasons covered above.
What this does not explain
None of the three terms, on its own or combined with the others, says anything about how large each correction should actually be, since knowing that proportional control needs strengthening does not say by how much, and turning any of the three terms up too far causes its own new problems rather than simply improving the correction further. Naming what each term is trying to fix is the easy half of PID control. Deciding exactly how much of each to apply to a specific real system, without overcorrecting into a new kind of instability, is the harder half, and it is where the actual work of tuning a controller, covered later in this set, really begins.
What follows from this
Once PID is understood as three separate fixes for three separate problems rather than one indivisible technique, tuning a controller stops being a matter of guessing at abstract numbers and becomes a matter of diagnosing which specific symptom, a lingering offset, a slow response, an overshoot, is actually present, and adjusting the one term responsible for it. This reframing carries directly into the rest of this set, where a PID controller's individual terms, particularly the derivative term, turn out to cause their own distinct and sometimes serious problems once they are pushed too far in pursuit of a correction they were never quite meant to provide alone.