← Product Building

Definition-of-done driven development

Product Building

For most of my career, the craft was in the how. How the code works, and how to structure the thing so the next change is cheap. That knowledge still matters — but it’s no longer where the leverage is. AI agents now produce the how faster than I can type it, and they’re improving quarter over quarter while my typing speed is not.

What agents cannot produce is the what. What does this product do for the person using it? What are its components, and what does each one owe the others? And the question underneath both, the one this whole approach hangs on: what, precisely, would have to be true for this work to be done?

I’ve started calling the practice definition-of-done driven development, because the name says where the effort goes. Before any code exists, you write the definition of done — a list of checkable outcomes, not activities. Then you design the verification that can confirm each one automatically. Only then does implementation start, and by that point implementation is the least interesting part.

Outcomes, not activities

The failure mode of most task descriptions is that they describe motion instead of arrival. “Add caching to the profile endpoint” is an activity. You can do it badly and still, technically, have done it. Compare: “the profile endpoint returns in under 200ms on a warm cache, returns correct data after a profile edit, and serves stale data for no longer than 60 seconds.” That’s a definition of done. Each clause is checkable by a machine, and — this is the part people miss — each clause is a product statement. It describes what the user experiences, not what the developer did.

Writing these is harder than it looks, and the difficulty is the point. Every ambiguity you resolve in the definition is an ambiguity the implementation can’t get wrong. Every clause you can’t figure out how to check is telling you something — usually that you don’t actually know what you want yet, which is much cheaper to discover before the code is written than after.

The test I use for each clause: could a person who didn’t write the code mark it done or not-done, from the outside, without asking anyone? If not, it’s not a definition. It’s a hope, and a hope has no finish line: work without one never gets ready, it only gets interrupted.

The verification system is the product decision

The thing that took me longest to see is that the tests are the design, stated precisely, rather than a chore that follows it.

When you design a product as components, the real design content is in the contracts — what each component promises the others, and what the whole promises the user. A verification system is just those contracts made executable. Deciding what to verify is deciding what the product guarantees; deciding what not to verify is deciding what it doesn’t, and that is product work of the highest order, exactly the work that doesn’t delegate.

So the job becomes: define the components, define what each owes, and build the harness that checks those promises continuously. I’ve written about why an agent’s usefulness is bounded by what it can verify about its own work — this is the same idea driven to its conclusion. If the loop is the constraint, then engineering the loop is the job, and the code inside the loop is a detail. Some people call this loop engineering. Definition-of-done driven development is the same discipline, entered from the product side: the loop exists to check a definition, so the definition comes first.

What this looks like in practice

A concrete shape, for a feature of ordinary size:

First, the definition. A short file, committed with the work: a one-line goal — who this is for and what they get — and a list of done-clauses. Outcomes a reader can check off, written before implementation starts. If a clause needs a diagram of components to make sense, draw it now; you’re deciding the architecture anyway, just at the level where it’s cheap to change.

Second, the checks. For each clause, decide how a machine confirms it: an end-to-end test that walks the user’s actual path, or an assertion against a real database. Build the missing harness pieces first — the fixture, and the one command that runs everything. This step is where the effort goes, and it should feel like most of the work, because it is most of the work.

Third, the implementation — through the loop. Hand the definition and the checks to an agent, or write it yourself with the checks running continuously. Either way, the checks are the arbiter. “It looks right to me” is no longer a state the work can be in; it’s either passing the definition or it isn’t.

Notice what happened to code review in this world. Reviewing the how becomes lighter — the checks carry most of that weight. Reviewing the definition becomes the serious act, because a wrong definition now gets implemented flawlessly.

The uncomfortable parts

Honesty requires listing where this bites.

A definition can be wrong. Verified-done and actually-valuable are different things, and no harness closes that gap. You can meet every clause and discover the feature was the wrong feature. The practice concentrates product judgment, which also means bad judgment concentrates. This is why keeping your own judgment sharp while delegating output stops being self-improvement advice and becomes an operational requirement.

Some qualities resist clauses. “Feels fast”, “doesn’t annoy people” — real requirements, hard to mechanise. The answer is to get them as close to checkable as honesty allows (budgets for load time, a walkthrough step for feel) and to be explicit that a human signs off the remainder. An unstated quality bar is the one that silently disappears.

The up-front cost is real. Writing definitions and building harnesses is slower on day one than just starting to type. The payback comes from every subsequent change to the same surface — the harness is already there, and “did I break anything” is a question that answers itself in a minute. It’s the same compounding argument as routing recurring work through skills: pay once in structure, collect on every run after.

Why product builders should care most

If you build products — rather than components to someone else’s spec — this shift is in your favour, and it’s worth saying plainly.

The scarce skill used to be split: the people who knew what to build were rarely the people who could build it, and the translation between them was where products went to die. Agents collapse the translation. A person who can define the product precisely — components, contracts, done-clauses, checks — can now get it built without owning every line of the how. The bottleneck moves from can you implement it to can you specify it so that done is checkable.

That second skill is rarer than the first ever was. Most people, asked what done means for their current project, will give you an activity list or a feeling. Learning to answer with a definition — sharp enough that a machine could referee it — is, I think, the highest-leverage skill a product builder can practise right now. The code was never the product. Now the tools agree.

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes