← General

Try the cheap fix first — unless failing it costs you the expensive one

General

Something is broken and there are two ways to fix it. One is cheap and might not work. The other is heavy and almost certainly will.

There are two pieces of standard advice here and they contradict each other. Do it properly the first time says go straight to the heavy fix, because you will end up there anyway and the detour is waste. Try the simple thing first says the opposite, because the cheap fix works often enough to be worth a go and you can always escalate.

Both are right sometimes, which means neither is a rule. What decides between them is what happens to the expensive option after the cheap one fails.

The obvious half of the rule

Order your attempts by what each one costs divided by how likely it is to actually end the problem.

That is genuinely most of it, and it explains why try the simple thing first is usually decent advice. A cheap attempt with a fair chance of working is a good trade: you pay a little for a real chance of paying nothing more.

It also explains the case people get wrong in the other direction. When the cheap fix has a poor hit rate and the expensive one is near-certain, trying cheap first is choosing to pay twice for the feeling of being careful. Restarting the service does not fix a schema problem, and everyone in the room knew that before anyone typed the command.

So far, so ordinary. The interesting part is that this arithmetic quietly assumes something that is often false.

The half that actually decides it

The arithmetic above assumes the expensive fix is still sitting there, at the same price, after the cheap one fails. That assumption does enormous work, and I have almost never seen anyone check it.

A cheap attempt is only cheap if failing it leaves the expensive attempt exactly as available, and exactly as cheap, as it was before.

That is the test, and it is entirely about what the attempt leaves behind.

Two things that look identical from any distance:

Everyone knows the difference in the abstract. Almost nobody applies it under pressure, because both start the same way: with somebody saying “let’s just try the quick thing”.

The tell, in software

Probes, all of which vanish completely if they fail:

Patches wearing a probe’s clothes:

The tell is consistent. A probe you undo by deleting it. A patch you undo by remembering it. Anything that relies on somebody remembering, later, under different pressure, with the incident closed, is not reversible in any sense that should count toward the decision.

What this cost me on a concrete floor

The slab under my tiling came out three centimetres off the flatness I had specified. Two shapes of answer were available. Level the slab, which is certain and expensive. Or change the bedding material to cement mortar, whose thickness varies enough to absorb the error inside the bed, which is cheap and would have ended the problem outright.

The second option was gone before the problem appeared, because the porcelain tile I had already chosen only bonds properly to thin-bed adhesive, and thin-bed adhesive has millimetres of range. So the cheap answer had been closed off by a decision made weeks earlier for reasons that had nothing to do with slabs. That is a whole argument of its own, and the short version is that low-tolerance components delete your cheap options before you know you need them.

The lesson that belongs here is narrower and it is about sequencing. The cheap option was cheap because failing at it would have left me exactly where I started, and that is the only property that ever mattered.

When the expensive fix is the one-way door

There is a third input, and it inverts the advice rather than adjusting it.

Everything above assumes the heavy fix is merely expensive. Sometimes it is expensive and irreversible: a destructive migration, a schema change other teams build against, an email that goes out, concrete that gets poured. When that is true, a cheap probe earns its place even at a poor hit rate, because a small chance of never having to open the one-way door is worth more than the probe costs.

This is the mirror image of ranking work by what cannot be undone rather than by effort. There, irreversibility decides what you do first. Here, it decides how much you are willing to spend to avoid doing it at all.

The order that falls out

  1. Is there a genuine probe available, one that vanishes if it fails and has a real chance of ending it? Run it first.
  2. Is the cheap option a patch rather than a probe? Then it is not cheap. Price it as the expensive fix plus interest, and it almost always loses.
  3. Is it a real probe, but with a poor hit rate, against an expensive fix that is certain and reversible? Go heavy immediately. Being thorough is not the same as being slow to admit what you already know.
  4. Is the expensive fix irreversible? Then probe even at a poor hit rate. You are buying a chance of not committing.

Set the probe budget before you start

The failure mode I keep seeing is trying the cheap thing eleven times.

Cheap attempts have a property that makes them dangerous in aggregate: each one is individually defensible. None of them is the moment you clearly should have stopped, and the eleventh is argued for by exactly the same reasoning as the first — except now it is being argued by somebody who has spent four hours and would like them to have been worth something. That is sunk cost arriving in miniature, hour by hour, where it is much harder to see than in the decade-sized version.

So name the number in advance. Two probes, or thirty minutes, and then we do the heavy fix. Written down before the first attempt, because the person deciding at attempt five is not neutral about attempts one through four.

Some problems will not show you their shape until you are inside them, and for those, starting is the only way to find out. Probing is how you start without paying for the privilege. The budget is how you stop.

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes