7/18/2026 · Ben Weller

What Not to Fix

What Not to Fix

An audit is built to find what is broken. Point trained attention at an operation and it will return a list of drag: the manual step, the doubled entry, the report someone rebuilds by hand every Monday. That part is close to automatic.

What to do with each line on the list is a separate question, and a harder one. The most disciplined line in an audit is often the one that names a piece of drag and says to leave it exactly where it is, a different call than finding the drag in the first place.

Some drag is cheaper than the fix

Finding an inefficiency and removing an inefficiency are two decisions, and only the first one is free. Spotting the doubled entry costs nothing but attention. Removing it costs whatever the fix costs, and that number is almost never the one written on the whiteboard.

The two get confused because drag announces itself while the cost of the fix stays quiet. A step that takes an extra thirty seconds is visible every single day. The afternoon it takes to wire two systems together, the week the new process takes to settle, the one case a year where the automated version does the wrong thing without anyone watching, none of that shows up in the moment you notice the thirty seconds. So the drag reads as pure loss and the fix reads as pure gain, and the arithmetic that would tell you otherwise never gets done.

Do the arithmetic and some drag turns out to be the cheaper side of the trade. The inefficiency is not good. It is only cheaper than its own replacement, and it stays the right answer until that changes.

A fix is never free

Take a distributor that keys every order twice. Once into the accounting system when the order comes in, and again onto the warehouse pick sheet before it goes to the floor. On any efficiency pass this is the first thing flagged: the same numbers, entered by the same kind of person, into two places that could obviously be connected. Wire accounting to the warehouse and the second keying disappears. Thirty seconds an order, gone.

Except the second keying was doing something nobody wrote down. When the order hits the pick sheet, it passes through a person who knows the floor, and that person catches the quantity that is off by a decimal, the SKU that was discontinued last month, the address that cannot take a full truck. The thirty seconds was never only data entry. It was a verification step wearing the costume of a clerical one. Connect the two systems cleanly and every error that used to die on her desk now travels straight to the floor at the speed of the new integration.

This is the cost that stays hidden until the fix exposes it. A workaround that has survived a while is usually holding more than one thing in place, and the load nobody documented is the load that breaks when the workaround comes out. That is the third cost of any fix, and the one most likely to be missed. The first two are easier to see and still get skipped: the work of making the change, and the risk that the change is wrong. Every fix carries all three. The build itself. The chance the new version has a flaw the old version did not. And whatever second job the old version was quietly doing while everyone assumed it did only the obvious one.

None of these three goes to zero as tools get cheaper or better. A future where building the integration takes an afternoon instead of a week still leaves the risk that it is wrong, and still leaves the buried verification step that the integration was never told to preserve. Cheaper building lowers one of the three costs. It does not touch the other two. That is why some drag stays worth keeping even as the tools improve. The tools were never the whole cost.

The keep decision expires

Keeping a workaround is a real decision, and like any real decision it can go stale. The thing to get right is what makes it go stale, because that is what tells you when to look again.

It is not the tools. A kept workaround does not become worth replacing because a better piece of software shipped. It becomes worth replacing when the business underneath it moves. Volume is the common one: the thirty-second check that was trivial at forty orders a day is an hour of someone's morning at four hundred, and the trade that favored keeping it flips. Dependence is another: something downstream quietly starts relying on the workaround, and now it holds up more than it was built to. And the quietest one is the person. The verification step on the pick sheet is only worth keeping while the person doing it can spot the bad SKU. The day she leaves, the check leaves with her, and a workaround that was correct on Friday is a liability on Monday, unchanged, because the thing that made it safe walked out the door.

Every one of those triggers is a fact about the business, not about technology. That is what makes them durable tests. Volume doubling, a dependency forming, an owner leaving, these will still be the reasons a kept decision expires long after the current tools are gone. So the keep is never permanent. It is correct for now, tagged with the specific condition that will end it.

This is worth separating from the other kind of stopping point, the one that does not move. Some boundaries in an operation are set by who is answerable when something goes wrong, and those hold no matter what changes around them. The keep decision is a softer thing. It is a bet that today's trade favors leaving the drag alone, and bets get re-priced when the odds move.

Keeping and fixing are both decisions

The point is to treat leaving something alone as a decision carrying the same weight as changing it, made the same way, with the full cost of the fix on the table instead of only the visible cost of the drag.

So the test at the end of any audit is two-sided. For each piece of drag, the honest cost of living with it, measured across a year and not a single morning. Against the full cost of the fix, all three parts, the build and the risk and the buried job. Where the drag costs more, fix it. Where the fix costs more, leave it, and write down the condition that will change the answer, so the decision gets looked at again on purpose instead of by accident when something breaks.

An operation that has done this ends up with a shorter fix list than one that has not, and a truer one. Everything left on it is there because someone weighed it and chose it. Everything taken off was taken off for a reason that will still hold in ten years. Spotting what looks broken is the easy half, and almost anything looks broken when the only thing you measure is the drag. Knowing what to leave standing is the half that takes the judgment.


No comments here. Ben reads email: ben@deployedskills.com.

All field notes