I specified a concrete slab that was flat and finished well enough to take an epoxy floor directly. What I got had three centimetres of variance across it.
That error was going to cost me something whatever I did next. What decided how much was a choice I had already made on a different day, about a different trade, for reasons that had nothing to do with the slab: I had picked porcelain tiles.
The chain, stated plainly
Porcelain is dense and barely porous, which is most of the reason people want it and the reason it will not bond well to a traditional thick bed of cement mortar. Porcelain wants thin-bed adhesive. And thin-bed adhesive, as the name concedes, is thin: its working range is millimetres where mortar’s is centimetres, so it can only go down on a surface that is already close to flat.
Cement mortar is the opposite kind of material. It goes down thick, and that thickness can vary across a floor. A tiler working in mortar absorbs an uneven slab inside the bed, adjusting depth as they go, and the finished floor comes out flat regardless of what the slab underneath it was doing.
So the sequence ran:
- The slab came out three centimetres off what I had asked for.
- The tile I had already chosen ruled out the material that would have swallowed that error.
- The error therefore had to come out of the slab before tiling could start: more material, more days.
Hundreds of dollars, and every one of them traceable to a decision that on the day looked like it was only about what the floor should look like. Which repairs are worth attempting before you commit to the heavy one is its own decision and has its own test.
Tolerance is a budget, and somebody always pays it
Every component in an assembly has a tolerance: the range of upstream variation it accepts without failing. That range is a budget, and it gets spent whether or not anybody planned for it.
A high-tolerance component absorbs variation inside itself. Cement mortar has centimetres of it, and it spends them on your behalf, quietly, without anyone having a conversation about it.
A low-tolerance component has nothing to absorb with, so it exports the variation to whatever it touches. Thin-bed adhesive requires a flat slab. Choosing it relocated the tolerance requirement — upstream, onto a layer that had already been poured and could no longer be argued with.
Precision in a component is not a property of that component. It is a demand it places on everything beneath it.
Where software does the same thing
Software does this constantly and almost never names it, because nothing arrives labelled this only works if the thing underneath you is exact.
- A strict parser. A validator that rejects an unknown field does not tolerate a producer that adds one. It has quietly converted every future change on the producer’s side into a coordinated release.
- A pinned toolchain. An exact version lock is the flattest possible substrate requirement. It buys reproducibility and charges you a broken build every time a machine, a base image or a runner drifts.
- Anything that trusts a clock. Ordering events by timestamp works while the clocks agree. It has no working range at all for the case where they do not.
- Exactly-once processing. A consumer that cannot survive seeing the same message twice demands a delivery guarantee nobody actually has. The forgiving version — make the handler idempotent — has metres of tolerance and needs nothing from the network.
- A fixed layout. A component sized to the content you had in front of you demands content of that size forever. A longer name, a translated string, a larger font, and it fails visibly, in front of the person you built it for.
In each pair, the forgiving option is the one that still works when something upstream — a producer, a runner, a clock, a translator, a contractor — does something slightly different from what you assumed it would.
The mistake was not the slab
This is the part I got wrong, and it is worth being exact about, because the obvious reading is that a contractor made an error and the error cost money. True, and not useful, because the contractor is not the part of this I control.
Here is the part I control. I chose a low-tolerance component that silently imposed a specification on a layer I had already been let down on.
The slab spec and the tile choice were made at different times, under different logic, and nothing in the process forced anyone to put them side by side. The tile decision presented itself as a question about appearance and durability. It was also, invisibly, a decision to require flatness from a slab that had already been poured and could no longer be negotiated with.
That failure has a general shape. Every time you adopt a component you inherit its preconditions, and its preconditions land on systems you may not own, may not have looked at in months, and may already have gone wrong. The adoption decision feels local. The precondition never is.
The software version is exhaustingly familiar once you start looking for it:
- A library that requires a newer runtime than the fleet is actually running
- A deploy strategy that assumes migrations are backwards compatible, adopted without checking whether this one is
- A cache whose only invalidation path is a write, in a system where something else also changes the data
- A pipeline that assumes every producer emits well-formed UTF-8
Each of those hands a bill to a layer that was not in the room when it was made.
Then why did I want a perfect slab?
Because none of the above is an argument for never being precise, and the line between the two cases is the only part of this worth memorising.
Errors at the foundation compound. Errors above it add.
A bad slab is a problem for the tiler, then for the skirting, then for door clearances, then for every threshold where this floor meets the next one — one error inherited by every layer sitting on it, arriving as a separate bill each time. Nothing above the slab behaves that way. A tile laid badly is a tile. You lift it and replace it, and the rest of the floor never hears about it.
Which gives a rule that survives leaving the building trade.
Spend precision where the error compounds. Buy tolerance everywhere else.
The foundation is any layer whose error is inherited by everything downstream, whatever its position in the stack, and in most systems there are only a few:
- The data model, because every query, report and migration is built on top of it
- Identity and permissions, because every feature reads them and none of them can safely disagree
- Public interfaces, because you cannot correct them without also correcting other people
- Naming, quietly, because it is the model everyone else reasons with
Above those, the same effort spent on precision buys much less, and it buys it at the cost of the thing you actually wanted, which is to be able to change your mind cheaply. This is the same instinct as ranking work by what cannot be undone rather than by effort: the compounding layers are the load-bearing ones, and they earn a different process, not merely a higher priority.
The check that would have caught it
One question, asked before adopting anything:
What does this choice assume about the layer underneath it, and when did anyone last verify that assumption?
Two clauses, and it is the second one people skip. Naming the assumption is usually easy once you look. Whether it is still true is the part nobody checks, because it was true when the layer was designed, and a design document is the last place anyone goes hunting for drift.
I have written before about designing my house the way I run production systems, and this is the rule I did not have then. Had I asked the question while choosing the tile, the answer would have been: this assumes the slab is flat to within millimetres, and I have not checked whether it is, which is a five-minute measurement. It would either have saved the money outright or moved the decision to cement mortar, where the slab’s error would have cost me nothing at all.