← AI Leverage

probe-ux: user testing with an agent that plays your users

AI Leverage

You cannot un-know your own product. You built the navigation, so you can’t get lost in it; you named the button, so it’s obvious to you. The product you shipped and the product a stranger meets are two different products, and you are structurally incapable of meeting the second one.

The traditional fixes are real user testing — the gold standard, and expensive enough in recruitment and scheduling that most teams do it rarely or never — and hallway testing, which is cheap but samples people suspiciously like yourself. So in practice most products are tested by the one population guaranteed not to find the problems: the people who made them.

I built a small agent skill called probe-ux to attack the gap, and after using it for a while I think the interesting thing about it is how little it insists on. I’m publishing it as a public repository on my GitHub so you can read the whole thing rather than take my word for it.

What it is

probe-ux is a skill — a written procedure an AI agent follows, in the sense I’ve described in why I route recurring tasks through skills. Point it at a URL, hand it a persona and a user story, and the agent works the product as that person attempting that story, in a real browser, clicking what a person would click and reading what a person would read.

Two inputs, and both are deliberately yours to supply:

The persona is whoever you need it to be: a retired teacher who distrusts forms that ask too much, say, or a first-time freelancer who doesn’t know the domain vocabulary your interface assumes — both invented here, to show the range. The skill doesn’t ship a canned persona library, because canned personas are how user testing turns into theatre. The point is to test with your users’ constraints.

The user story is the journey you care about: can they get from landing to signed up, can they find the pricing basis before committing, can they complete the headline task and know they’ve completed it, would anything bring them back tomorrow. The skill walks the story and reports where it stalled, what was confusing at each step, and — the part I find most valuable — what the persona believed was happening at moments where the interface left it ambiguous.

That’s the whole design. It’s un-opinionated on purpose: no scoring rubric baked in, no framework it’s married to. Give it a different persona and story and it’s a different test. The skill supplies the discipline of the walk; each run supplies its own judgment about what matters.

Why the output is different from a checklist audit

There’s no shortage of automated UX audit tools, and most produce the same artifact: a list of violations against someone’s heuristics. Contrast ratios, missing labels. Useful, worth running, and still a different thing from user testing. A page can pass every heuristic and still lose the user, because heuristics check surfaces and users experience journeys.

What a persona walk produces instead is a narrative with failure points. From an invented run, to show the shape: At this step, I didn’t know whether my payment had gone through, so I pressed the button again. I abandoned here because the form asked for company details and my persona is a sole trader — nothing said I could skip it. I finished the task but couldn’t tell I’d finished; the success state reads like another form. These are the findings that fill the gap between “technically works” and “a stranger can actually use it”, the kind you otherwise only get by watching a real person, on the day you can’t easily arrange one.

It’s cheap enough to run on every significant change rather than once a quarter. And because the walk is agent-driven, it composes the way anything skill-shaped composes: run the same story as three different personas in parallel, or the same persona across your product and the two competitors a real buyer would also open. It will happily walk products you don’t own, which is worth doing at least once, if only to calibrate how your funnel actually compares to the ones your customers came from.

Read-only, and report-only, on purpose

Two constraints in the design carry most of my opinions about how this class of tool should behave.

It’s read-only. The skill navigates and observes; it creates no accounts with real consequences and mutates nothing it can avoid mutating. Partly safety, partly honesty: the moment your test harness starts changing state, your findings are entangled with side effects, and running it against a production system — or a competitor’s — stops being defensible.

It reports; it never prescribes fixes. The output is what a stranger experienced, not a redesign. That boundary matters more than it looks. The moment a tool starts proposing solutions, its findings bend toward the solutions it likes, and deciding what to fix and in what order is product judgment that shouldn’t be delegated to the instrument that gathered the evidence. I’ve written about why handing that layer over is the expensive mistake; an agent that impersonates your users is precisely where the temptation to also let it do your thinking is strongest.

What it can’t do, stated plainly

An agent playing a persona is a simulation, and simulations flatter you if you forget what they are. It models where frustration is likely rather than feeling any, which is weaker evidence. It won’t reproduce the real-world texture of a distracted user on a cracked phone screen on a train. And it carries whatever assumptions the underlying model has about how people behave, which are averages, and your users may not be.

So the honest position for it: far cheaper than real users, and no replacement for eventually watching an actual human. Where it earns its keep is frequency — catching the obvious stumbles continuously, so that when you do spend real-user time, it’s spent on the subtle findings only real users can give you rather than on discovering that nobody can find the login button.

The deeper pattern is the one I keep returning to: once a practice is encoded as a skill, it stops depending on anyone remembering to do it, and starts compounding. User testing has always been the discipline everyone endorses and nobody schedules. Making it a thing an agent does on request moves it into the category of things that actually happen.

Hsein Bitar is a product developer, DevOps and backend engineer. He owns infrastructure, CI/CD and backend architecture in production at NSquared. Who this is, and what he ships.

Read more notes