← Back to Archive

What a fuel gauge has in common with a month of software

Why fitting a small sensor is usually cheaper than writing the code that would avoid it.

A fuel gauge and a month of software are two answers to the same question, which is how much is left, and the float in the tank wins not because it is cleverer but because it is in contact with the quantity being asked about, whereas software that works the same answer out has to assemble it from other measurements that are merely related to it and inherits the error sitting inside every one of them.

I spent a fortnight on an estimator for a state that a switch would have reported directly, on the grounds that the switch was another part to seal, wire and eventually replace, and by the time the estimator was behaving well enough to trust in the conditions we could reproduce, the fortnight had cost more than fitting that switch to every machine we expected to build.

The mechanism behind an estimated measurement

Software that infers a physical state almost always does it the same way, by starting from a moment when the state was known and then adding up everything that has happened since, which is dead reckoning whatever else it is called in the code. A wheel turned this many times, so the machine moved this far; the motor drew this much current for this long, so the mechanism must have reached the end of its travel; nothing has changed on the input, so the door must still be where it was left. Each of those steps is individually reasonable and individually small in its error, and the trouble is that the errors do not cancel, because they are not random with respect to the thing causing them. A wheel that slips on wet ground slips in the same direction every time, so the estimate does not wander either side of the truth, it walks steadily away from it.

What a direct sensor changes is not the accuracy of any single reading but the length of time over which error is allowed to accumulate. An estimator's error grows with however long it has been since the last moment of certainty, whereas a sensor makes that interval equal to the gap between one reading and the next, which on most machines is a few milliseconds. That is the entire difference, and it is why an estimator that performs beautifully over a short run can be hopeless over a long one while a cheap switch is equally right in both cases.

The fuel gauge comparison

The range estimate on a car's trip computer is built out of two things measured rather well, namely distance travelled and fuel consumed over the recent past, and it is still wrong on a cold morning with a full car and a hill ahead, because underneath the arithmetic it is assuming that the next hundred miles will resemble the last hundred. The float in the tank assumes nothing at all. It sits in the liquid and moves when the liquid moves, and a fuel gauge built around it is right on the hill, right with the trailer, and right after the car has been parked on a slope for a week, none of which the trip computer can manage without being told about each situation in advance.

The reason this comparison is worth holding on to is that nobody would propose deleting the float and computing the fuel level instead, even though it could be done, and the objection is not that the computation is impossible but that the float already answers the question completely and costs almost nothing. The same proposal is made constantly about machines, in the form of an argument that a state can be worked out from information the system already has, and it is much harder to refuse there because the sensor has not been fitted yet and so its absence does not feel like a loss.

One figure worth keeping in mind

A limit switch with a magnet and a short length of cable is a component cost, roughly what a sandwich costs, while the software that exists to avoid fitting one is priced in engineer weeks, and the gap between those two numbers is three or four orders of magnitude before anyone has argued about anything. The argument nevertheless keeps happening, and it keeps happening because the two costs behave differently as the fleet grows, since the sensor is paid for again on every machine while the software is paid for once. That makes it a genuine crossover question rather than an obvious one, and the useful thing to notice is where the crossover actually falls, because a sandwich per machine against a fortnight of salaried time only breaks even somewhere in the tens of thousands of units, which is a scale most hardware projects never see and none of them see early.

Where this stops being true

There are real cases where the estimate is the right answer, and the clearest is a quantity that genuinely cannot be reached, either because there is no room for a sensor where the measurement would have to happen or because nothing that could survive the environment would fit. There is a second case worth being honest about, which is that every sensor added to a machine that lives outdoors brings a connector, a cable and a seal with it, and a cable that moves is a reliable way of introducing a fault that the estimator never had. The third case is the most dangerous, because a sensor that is quietly wrong is worse than an estimate known to be approximate, since everything downstream treats a measurement as ground truth and stops sanity checking it.

What this changes in practice

The question worth asking at the start of a problem like this is not how the state should be estimated but what the smallest physical change would be that makes the state trivial to read, because that question has an answer that can be costed in an afternoon while the other one cannot be costed at all until the code is nearly finished. It also changes what testing looks like, since every state that is estimated rather than measured has to be validated across every combination of conditions that could bias the estimate, and that set of combinations grows far faster than anyone budgets for, whereas a switch is tested by pressing it.

My key error with this

I had a rule that every added component is another thing to seal, wire, source and eventually replace, which is true, and I applied it to sensors as though they were the only components on the machine that carried that cost. What I had not seen is that an estimated state is a component too, with a purchase cost in engineering time, a recurring maintenance cost every time the machine or its environment changes, and a failure mode that is harder to diagnose than a broken wire, and the only real difference is that it never appears on the bill of materials and therefore never gets counted. Since then the first thing I look for is a physical change that deletes the estimate, and I have not once regretted spending a component to save a subsystem.

More on Finding the fault