← AI Leverage

Stop doing the task. Run it through an agent skill.

AI Leverage

Work you do by hand disappears when you finish it. Work you route through an agent skill — a written procedure an AI agent executes, with scripts for the parts that never change — accumulates. Both ways finish the task. The second also leaves the process better than it found it, and the gap between the two compounds every week you repeat the work.

On the first run the skill is slower, because you have to write the thing, so speed is the wrong frame. What differs is what each approach optimizes. Doing a task by hand optimizes the instance. Doing it through a skill optimizes the process, because the process now exists as an artifact you can read and criticise.

A process you can see is a process you can fix

Here’s the part that surprised me in practice. The biggest win wasn’t automation. It was exposure.

When a task lives in your head, you can’t inspect it. You do the steps in whatever order habit dictates, and the inefficiencies are invisible because nothing forces you to articulate them. The first time you write the task down as a skill — the trigger, the steps, the checks, the output — you’re forced to look at the whole shape of it at once. Every time I’ve done this, the writing alone surfaced something: a step that existed for a reason that no longer applies, two steps that could be one script, a manual lookup that a command could answer, a decision I was re-making every run that should have been made once.

And it keeps happening. Each run of the skill is a small review of the procedure. The agent hits an ambiguity, and you notice the instruction was vague. A step keeps failing in prose, and you realise it was a script waiting to be extracted. The task becomes a feedback loop on itself, which is a property hand-work never has. Doing a task twenty times by hand teaches you very little about the task, because nothing records what varied. Running it twenty times through a skill produces twenty diffs against a written baseline.

The skill is the documentation

Every team says they’ll document their processes. Almost none do, because documentation is a separate task that competes with real work and always loses — I’ve written elsewhere about pointing AI at exactly this category of chronically-skipped work.

A skill dissolves the problem. The procedure and its documentation are the same file. There is no drift between “what the runbook says” and “what we actually do”, because the runbook is what runs. When the process changes, the skill changes, because it has to — the old version stops working. Compare that with a wiki page, which keeps asserting last year’s process indefinitely, to anyone unlucky enough to trust it.

This matters most at handover. A task that lives in one person’s head is a task only that person can do, and the transfer cost is a training session nobody schedules. A task that lives in a skill transfers by pointing at the file.

Missed steps stop being a personality trait

Some people never skip steps. I am not one of them: not the same step every time, which would at least be diagnosable, but whichever step the day’s interruptions happened to land on. The classic fix is a checklist, and checklists genuinely work. Aviation runs on them.

A skill is a checklist that executes. The difference matters because a paper checklist relies on the discipline of the person holding it, which is exactly the resource that’s depleted on the days you skip steps. The skill doesn’t have days. Step four happens after step three every single run, including the run at 6pm on a Friday and the run where you were interrupted twice in the middle.

Delegation stops being all-or-nothing

Without a written procedure, delegating a task means delegating to someone who already knows how to do it, or accepting a long apprenticeship. With the task encoded as a skill, delegation becomes granular. You can hand the whole thing to an agent and review the output. You can hand it to a colleague who’s never done it, because the skill carries the knowledge. You can split it — the mechanical middle automated, the judgment calls at either end kept — and that split is visible in the file, so everyone can see exactly which parts are trusted to run alone.

That granularity also changes how load spreads. Work that only you can do queues behind you; you are the bottleneck and the single point of failure. Work encoded in a skill can run in parallel, or run three times against three inputs while you’re doing something else. The task stops being attached to a person and becomes attached to a definition.

Verification is a layer you add once

The quiet failure of most recurring work is that “done” is asserted rather than checked. The report was sent — was it right? The deploy went out — did it work? Checking is a whole second task, so it gets skipped even more reliably than documentation does.

Once a task runs through a skill, verification has somewhere to live. You add a check step at the end — a script that asserts the output looks like the output should look, a command that exercises the thing that was just changed — and now every future run is verified, forever, at no additional discipline cost. I’ve argued before that an agent is only as good as the feedback signals it can read; the same is true of a process, and a skill with a verification step is one that tells you when it broke.

This stacks with everything above: the verification is documented because it’s in the file, and it makes delegation safer, because the skill checks its own work regardless of who — or what — ran it.

Where to start, and where not to

The test for whether a task belongs in a skill: have you done it three times, and will you do it again? One-off work doesn’t repay the writing. Judgment-heavy work — the kind where every run is genuinely different — belongs mostly outside the skill, though its scaffolding often doesn’t. My probe-ux skill is deliberately built that way: the procedure is fixed, the judgment each run needs is supplied fresh.

Start with the task that annoys you most on repetition. Write down what you actually do, honestly, including the checks you’re supposed to run. Extract every step that’s the same twice into a script — a step that keeps failing as prose is a script you haven’t written yet. Then run the skill instead of the task, and edit the file every time reality disagrees with it.

The first run costs you an hour; every run after that pays interest. And once “done” is defined well enough for a skill to check it, you’ve started practising definition-of-done driven development, and that changes more than the task.

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes