← AI Leverage

An AI agent inside a customer's workspace

AI Leverage

An AI agent that runs inside a customer’s workspace is the most demanding version of shipping inside a customer’s environment. It has judgment. It reads the customer’s data, follows an operating playbook, does work, and reports back. When it goes wrong, it goes wrong with initiative, on data you do not own, in a run the customer paid for.

I have spent a good part of this year building and hardening exactly that: an agent dispatched into a workspace after an operation like a dashboard clone, running a written skill against the customer’s state, and returning an outcome through a callback. Here is what broke, in order of how much it taught me, and the rule each fix left behind. The rule at the top of all of them: the customer’s run is the worst place to discover anything.

The instruction that contradicted the playbook

The agent’s instructions told it to delete a cloned chart under a certain condition. The team’s own playbook for that operation says park, never delete. Two documents, one drifted copy, and the agent was reading the wrong one.

The fix was to delete the copy and make the playbook itself the skill the agent runs, so there is one source of truth and the agent executes it directly. I have argued before that a task should run through a skill because the skill becomes the documentation; inside a customer’s workspace that stops being a productivity argument and becomes a safety one. If humans and agents read different procedures, the agent will eventually do the thing the humans agreed never to do.

The run that lost its own answer

A long agent run finished. The work was done. Then the callback carrying the result failed to connect, and the outcome was simply gone. The compute was spent, the customer’s data was touched, and nothing recorded what happened.

Then a second variant: the run outlasted its callback token, finished, and was refused with a permission error on the way back. Then a third: a shared, unkeyed rate-limit bucket plus a payload fetch that did not retry killed dispatched runs before they started, silently.

Each got the same shape of fix. The result path retries. A service-side backstop records the outcome even if the callback never lands. Tokens live as long as the run can. Each fix arrived with a test that reproduces the silence first, goes red, and only goes green when the outcome survives, which is the principle that runs through the whole series: every silent path gets a loud one.

The dependency that only failed in production

Agent skills are scripts. Scripts import libraries. The runner image is built from a pin list, and a skill can import something the image does not have. Nothing in the pull request notices. The failure surfaces the first time the skill runs, which is inside a customer’s autonomous run.

The fix was small and moved the failure a long way: the pins were made complete, and a CI test now fails the pull request if a skill imports anything the runner lacks. The rule is where the discovery happens. A dependency error at review time costs a minute. The same error in a customer’s run costs the run and the trust.

The boot step that deleted the tool

A local end-to-end loop for the agent container, with developer-only dispatch seams, existed mainly so that the whole path could be run without a customer in it. Its first real run caught a cleanup step that removed the pinned CLI during boot. Every production run would have failed at start.

Nothing about that bug was subtle. It was only invisible because there had been no way to run the path locally. The loop earned its keep on the first pass, which I have come to read as the sign that it should have existed earlier. An agent with excellent judgment and a broken callback is a very expensive way to do nothing.

I wrote about not letting AI erode your judgment as a personal discipline; inside a customer’s workspace it is an engineering one.

Where will your agent’s next failure be discovered: in a loop you run, or in a run the customer paid for?

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes