← AI Leverage

Clever code got cheap. Clever decisions did not.

AI Leverage

I used to guard every line that entered my projects. Most of that guarding has stopped paying for itself, and the cleverness it protected has moved somewhere else.

For a long time I held one bar. A change got into a project I owned only if I understood every line of it, the person who wrote it could defend each choice, and the next person could change it without fear. A change that worked but was hard to read went back. I was proud of that bar, and I kept holding it for a while after it stopped earning its keep.

What the bar was really buying

With Opus 4.6 and the models that followed, something moved underneath it.

Maintainable code was never a virtue on its own. It was insurance. Code gets changed later, changing code nobody can read is slow and risky, so you paid up front, in review time and rework, to keep that later change cheap.

The models collapsed the price of the later change. Reading an unfamiliar module, explaining what it does, refactoring it, rewriting it from a clear description: each of those went from an afternoon to minutes. When the thing the insurance covers gets that cheap, the premium stops making sense.

So on the backend my questions got shorter. The grand version of this is a new theory of code quality. The ordinary version is three checks on a normal Tuesday:

Code that passes all three and could clearly be better goes in. It gets improved on the day improving it becomes someone’s job, and that day now costs very little.

The second check carries more weight than it looks. “It breaks nothing” is only a claim if something can check it, which is why the feedback loop decides what an agent is worth. Without that loop, relaxing the bar is just hoping.

Where the cleverness went

Clever solutions still cost what they always cost: thinking, experience, and time spent inside the problem. What changed is where they pay back.

In the backend, a clever solution now saves lines a model would have written anyway, readably enough. In front of a customer, a clever solution is the difference between a screen someone understands on first sight and one they need a call to get through. Nobody rewrites that for free. When it is wrong, the cost lands on a person, and it keeps landing every day until someone notices.

That is where my care goes now: which default a form opens with, what an empty page says to someone who just arrived, which of three buttons should exist at all. I have argued that software should guide people to the outcome rather than present them a menu, and those decisions are what guiding is made of. A model will build whichever version you describe. Deciding which version deserves to exist is still judgment, and it is still yours.

Spend the cleverness where a rewrite does not undo the damage: in front of the customer, and at the security boundary.

Where I might be wrong

The evidence does not all point my way, and some of it points hard the other way.

GitClear looked at 211 million lines of code and found that blocks of five or more duplicated lines grew eightfold during 2024, while refactoring fell. That is the “works, could be better” code I now let in, measured at scale. My answer is that duplication hurts when a person has to change it by hand, and I am betting that the one changing it is increasingly a model. That is a bet, and I hold it as one.

METR ran a controlled trial with experienced developers working in their own open-source repositories and found that with AI tools they took 19% longer, while believing they had been faster. If the rewrite is not actually cheap, the insurance still pays, and my argument loses its floor. The study used early-2025 tools, and the models since then are the ones I am describing, but I would rather name it than wave it away.

Veracode found that AI-generated code introduced risky security flaws in 45% of its tests, and that larger, newer models did no better. This one I accept in full. It is why the third check exists, and why a security boundary is the one place I still read every line myself.

On the other side, Andrej Karpathy named the far end of this “vibe coding”: accept every change and forget the code exists. I stop well short of that; a change nothing can verify does not go in, whoever wrote it. Closer to where I stand are Jason Liu, who puts it as “ability is basically free… you don’t know what’s worth making”, and Matt McGuire’s case that code will be cheap and judgment will not.

So I would narrow my own claim. The tolerance covers the lines inside a function. It stops at the shape of the data, the boundaries between services, and anything a customer or an attacker touches, because those stay expensive to change whoever does the typing.

A few years out

A few years from now, most of the backend code merged this month will have been rewritten by a model at least once, and nobody will remember how it read. The first screen a customer saw will be remembered by the customer.

Code you can rewrite in minutes deserves minutes; a decision a customer lives with deserves the thinking. Where did your last hour of careful thought go: into lines a model could rewrite, or into a screen a customer has to live with?

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes