“Queued” can be true and useless at the same time. A partner operator sets up a new client and asks for the partner’s dashboards to be cloned into the client’s workspace. The system accepts the request, answers “queued”, and the operator moves on. A few minutes later the clone fails: the client’s workspace does not yet have the source connections the dashboards read from. The failure is correct; the timing, I think, is the defect.
The operator now has to find the failure, work out what it means, fix the connections, and ask again. Every one of those steps happened after they were told the work was underway, by an answer that confirmed the request without checking whether the request could succeed.
This is one spoke of a series on shipping inside a customer’s environment. The rule it argues for is short: check readiness at the moment you promise, not after.
What the pre-flight does
The fix was not to make the clone more tolerant. It was to move the check in front of the confirmation. Before anything is queued, the system resolves what the clone will need in the destination and compares it against what the destination has. If something is missing, the operation is held closed and the answer names each missing item, so the operator’s next action is obvious and immediate. If nothing is missing, “queued” now means “will succeed on the preconditions we can see”.
Three properties made it worth doing properly rather than as a quick guard:
- It names things. “Missing: two source connections” is a task. “Failed” is a search.
- It holds the operation closed. A warning that still lets you proceed is a warning people click through. The gate is the point.
- It answers at the moment of asking. The whole cost of the old flow was the gap between the promise and the truth. The pre-flight closes the gap to zero.
The bug the pre-flight nearly shipped with
Checks have scope, and scope is where readiness checks go wrong. The first version of this pre-flight resolved the destination’s connections without scoping the lookup to the destination workspace. On a single-tenant test it worked. Across real tenants it would have reported every onboarded client as missing everything, because it was comparing the partner’s needs against a set that was not the client’s.
That was caught in review, before it reached anyone, by asking the question every readiness check has to answer: “whose state am I reading?” I think a check that reads the wrong tenant’s state is worse than no check, because it fails confidently.
Where else this applies
The pattern applies anywhere a system confirms a request whose success depends on state the requester cannot see:
- A deploy that accepts and later fails on a missing secret.
- An import that starts and dies on a schema mismatch it could have diffed first.
In each case the fix I would reach for has the same shape: enumerate what the work will need and check it in the request path, then refuse with names or confirm with confidence. Whatever you cannot check before answering, say so in the answer.
If you have read my piece on definition-of-done-driven development, this is the same idea applied to a single request: “queued” is a claim about done, and it should be one you can stand behind.