← AI Leverage

Why AI coding agents aren't making your team faster

AI Leverage

The promise attached to coding agents was a three to ten times increase in output. A lot of companies are now reporting something in the region of ten percent, and a gap this large usually means everyone is measuring the wrong thing.

Here’s my explanation, and it’s testable on your own repo this afternoon. I think the multiplier is real, but it’s a multiplier on a feedback loop. If the loop is missing, there’s nothing to multiply.

What a coding agent is actually doing

An agent works in a cycle: make a change, check whether the change is correct, adjust. Same cycle a person uses. The difference is speed: an agent can go round that loop far faster than a human can, provided each lap produces a signal.

Everything hinges on the check step, and it’s the step nobody audits.

A human developer who can’t run the tests still has judgment and a mental model built over months. They degrade gracefully. They slow down and go and ask someone. An agent has none of that. Take away the check step and it doesn’t slow down at all. It keeps going at full speed in an unknown direction, which is strictly worse than stopping.

The question that matters is this:

What can the model verify about its own work, in under a minute, without a human in the loop?

In most codebases the honest answer is: almost nothing. Which is also the answer to why your multiplier is ten percent.

Measure your own loop in five minutes

Before you change anything, get the number. Time this, on a real change, on a normal working day:

  1. Make a one-line change that should break something.
  2. Start a timer.
  3. Run whatever your team runs to find out.
  4. Stop the timer when you know, not when the run starts, not when someone says “looks fine to me.”

That elapsed time is your loop. It’s the single most predictive number about what an agent will be worth to you, and almost nobody has measured it.

Under a minute and fully automated, an agent compounds it, because it can afford to be wrong twenty times an hour and still finish ahead. Ten minutes and it’s still useful but the agent spends most of its time waiting. Hours, or it needs a person to look at a screen, and the agent is running unverified, and unverified cycles are worth roughly what those companies are reporting.

While you’re at it, note the second number: how many of your test failures last month were real bugs versus flakes or someone else’s broken branch. A loop that cries wolf is a loop an agent learns to route around, the same way your team already has.

The Stripe example, and why people read it backwards

Large-scale rewrite work at Stripe is the reference point people reach for, and usually for the wrong reason. The interesting part is that the codebase had the properties that made mechanical, verifiable transformation possible in the first place: strong typing, dense test coverage, and consistent internal patterns that turn “did this change break something” into a question a machine can answer on its own.

That precondition was built over years, long before anyone pointed an agent at it. Teams reading the outcome and expecting similar results on a codebase without those properties are reading the result and skipping the setup.

The uncomfortable version: the teams best positioned to get a large multiplier from AI are the ones who already invested heavily in engineering discipline. The tooling amplifies an existing gap rather than closing it. If you were hoping agents would let you skip the boring decade, they won’t. They’ll price it.

The failure mode nobody warns you about

Here’s the specific thing that makes this worse than a simple absence of speedup.

Language models are extremely good at satisfying the stated objective. Ask for passing tests and you’ll get passing tests, which is a narrower thing than working code.

I’ve watched agents write tests that mock the exact boundary the test was supposed to exercise. Tests that assert a function was called rather than that it did anything. Tests that construct an input which can’t occur in production, verify the handler for it, and go green. My favourite, because it’s so nearly reasonable: a test that catches the exception the code under test is supposed to prevent, and passes, because catching it counts as handling it.

None of that is the model behaving badly. It’s the model being clever, routing around an obstacle to reach the stated goal, which is exactly what it’s built to do. The obstacle happened to be the part that mattered.

The result is a suite that’s green and steadily less connected to whether the critical paths work. Coverage and confidence go up. Actual verification goes down. And because the signal is now actively misleading rather than merely absent, the team moves faster in the wrong direction than they would have with no tests at all. This is the same trap as any metric that ranks your options correctly while measuring the wrong thing: the number moves, everyone relaxes, and nobody checks what it’s made of.

What to build before expecting a multiplier

The work is unglamorous and it is the whole job.

Tests that fail for the right reason. The bar is that when a test fails, the failure names the broken behaviour. One test that can only fail when the code is genuinely wrong beats twenty that fail when an unrelated import moves.

One command that tells the truth. If verifying the system takes a sequence of steps and knowing which three tests are flaky, an agent can’t use it, and neither can a new hire, which is the tell that this was always a problem and agents just made it expensive.

Critical paths written down. Name the flows that must never break. Then make sure each has a test exercising the real path (real database, real boundaries), not a mocked cartoon of it. These are the tests that must not be allowed to go green through cleverness, and they’re worth reviewing by hand.

Assertions on behaviour, not on calls. “This function was invoked with these arguments” tests your implementation. “Given this input, the system produced this output, and this row exists afterwards” tests your behaviour. Only the second survives a refactor, and only the second catches the clever workaround.

Types where they constrain. Types are the cheapest continuous verification available. They run on every keystroke and they narrow the space of changes an agent can make without noticing it broke something.

Every item on that list is worth having whether or not you ever point an agent at anything, which is why this isn’t really an AI post.

The order that works

I’d instrument the loop first, then bring in the agent.

Doing it the other way (agent first, hoping the speed pays for the missing verification) produces exactly the outcome those companies are reporting, because a multiplier applied to a signal you don’t have returns a signal you still don’t have. Once the loop exists, it also changes what deserves your care: plain code that passes it can go in, and the cleverness moves to what the customer sees.

If you want the version of this that isn’t about speed at all, the more durable win is pointing these tools at the work you reliably never get round to rather than at the work you’re already good at.

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes