Mirror

The Two-Minute Check I Run Before Acting on a Machine's Answer

I sent a client a wrong number that a tool gave me confidently. Here is the check I built afterwards, and the category of question it will not save you from.

In March I put a figure in a client document that I had got from an AI tool, that was wrong, and that I did not check because it was the kind of thing I would not have doubted from a colleague. The correction cost me an hour and a certain amount of standing. This is what I do now, and I want to be honest about what it does not cover.

a confident printed answer on a desk with one number in it circled and a second document being held up beside it to compare

The check is two questions and it takes about two minutes.

First: does this answer contain a specific that I could look up? A number, a name, a date, a feature that either exists or does not. If yes, look up exactly one of them — the one the conclusion depends on most — at the source. Not a second AI query; the vendor’s own page, the documentation, the actual invoice. One lookup, because the failure mode is almost never that everything is wrong; it is that one load-bearing specific is confidently invented and the surrounding reasoning is fine.

Second: would I be able to tell if this were wrong? This is the question I skipped in March. If the answer is no — if I have no independent way to evaluate it and I am accepting it because it is fluent — then I should not use it in anything that leaves my desk. That does not mean the answer is worthless; it means it is a lead rather than a fact.

Those two questions catch, in my experience over about six months, the great majority of the errors that would have mattered. They cost two minutes on maybe one query in four, since most of what I ask does not produce a specific that anything depends on.

Now the part they do not cover, which I think is underdiscussed. They only work on errors of commission — something asserted that is false. They do nothing for errors of omission, where the answer is entirely true and leaves out the thing that would have changed your mind. I have been caught by this twice, both times on questions about whether a tool could do something. The answer described what it could do, accurately, and did not mention the licence tier required, which was the whole decision.

I do not have a good general solution to omission. What I do is narrower: for any question where I am deciding rather than learning, I ask what would have to be true for this to be the wrong choice, and I go looking for that specifically. It is slower and I only do it for decisions above a certain size.

The broader thing I would say is that the failure in March was not really about the tool. I have taken numbers from colleagues without checking them, and from search results, and from my own notes from six months earlier. The difference is that with all of those I have a calibrated sense of how often they are wrong and in what way. With a machine that is fluent about everything at the same level of confidence, that calibration does not form on its own, and until it does the check has to be explicit.