Mirror

I Logged Every Time I Used an AI Tool for a Month. Two Uses Were Worth It.

Forty-one logged uses, honestly timed against doing it myself. The wins were narrow, real, and not the ones I expected.

For a month I wrote down every time I reached for an AI tool at work, what I was trying to do, and roughly how long the whole exchange took including reading and fixing the output. Forty-one entries. My honest assessment is that eleven saved me meaningful time, six cost me time, and the remaining twenty-four were a wash that felt productive.

a tally sheet of forty-one marks where eleven are circled in green

The wash category is the interesting one and I want to start there, because it is where I think most people are. These were cases where I got a plausible answer quickly, felt efficient, and then spent as long verifying it as I would have spent doing it. Writing a short piece of copy is the clearest example: thirty seconds to generate, four minutes of adjusting it into something I would actually send, versus about five minutes to write from scratch. It is not worse. It is also not better, and it felt much better, which is a gap worth being suspicious of.

The eleven genuine wins fall into two shapes with almost no exceptions.

The first is transformation where I already know what correct looks like. Turning a messy pasted table into structured data, converting a list of dates into a different format, rewriting a paragraph in a shorter form when I have the original in front of me. In all of these I can verify the output at a glance, because I have the input, and the failure mode is visible rather than subtle. Six of my eleven wins were this.

The second is getting started on something I was avoiding. Not the finished work — the first bad version that makes the blank page go away. Three of my wins were this, and I think it is a real effect and not a trick. My own first draft of a difficult email takes twenty minutes because I keep stopping; a mediocre draft I can react to takes four minutes to fix. The value is not the text, it is that reacting is easier than initiating.

The six that cost me time were all the same failure: asking about something specific that I could not verify quickly. Details of a pricing structure, how a particular feature behaves, whether a specific integration exists. The answers were fluent, and two were wrong in ways I only found out later, and one of those wrong answers went into a document I sent to a client. Recovering from that took the better part of an hour and a slightly embarrassing correction.

So the rule I now use is a single question before reaching for it: can I check this answer faster than I could produce it? If yes, it is probably a win. If no, I am not saving time, I am moving the work to a place where the errors are harder to see.

Two smaller observations. The time saved is real but it is not concentrated — eleven wins over a month came to maybe ninety minutes total, spread in three- and five-minute pieces, which is not the kind of saving that produces a free afternoon. And my log almost certainly flatters the tools, because I only wrote down the times I chose to use one, and the choosing is itself informed by an expectation of success.

I am not arguing anyone should use these less. I use them daily and will continue. I am arguing that “it saves time” is a claim that can be measured, that measuring it took me about ten seconds per use, and that my estimate before measuring was roughly three times my estimate after.