The Month in Numbers
I started this blog four weeks ago with a baseline test and a paper notebook.
Since then, I've run about 20 tests on real work tasks. Not demos. Not hypotheticals. Things that needed to be done by Friday.
Here's what the notebook shows.
Total tests run: 21
Tools evaluated: 14 (some free, some paid, some trials)
Tools kept: 5
Tools cancelled or removed: 7
Tools still in testing: 2
Total time saved: about 2 hours per week
Total time spent testing and correcting: about 3 hours per week
That last number matters. For the first month, I spent more time testing than I saved. That's the cost of building a baseline. It won't stay that way, but it's the honest picture.
What I Kept
Five tools earned a place on my desk. Here they are, in order of how often I use them.
1. Claude (free tier) — for drafting and rewriting.
This is the tool I use most. It handles the widest range of tasks: emails, summaries, rewrites, comparisons. It was the best performer in the three-way email test and the vendor follow-up test. It's not perfect, but it requires the least correction.
2. Extension A (page summarizer) — for long documents.
This is the one browser extension that earned its place. It summarizes long articles, policies, and PDFs in seconds. It drops caveats, so I always check against the source. But the time saved is real.
3. A paid transcription tool — for meeting notes.
It handled background noise better than the two free tools I tested. Correction time was about 11 minutes for a 45-minute meeting, versus 16 and 22 for the free alternatives. That's worth the subscription.
4. A general-purpose assistant (free tier) — as a backup.
This is the tool I reach for when Claude is down or when I need a second opinion on a draft. It's not as strong as Claude for the tasks I care about, but it's reliable and free.
5. My paper notebook — for tracking tests.
This isn't an AI tool, but it's the one thing that made the rest of this month useful. Without the notebook, I'd be guessing at every number I've written here.
What I Rejected
Seven tools didn't earn their place. Here's why.
Extension B (rewriter). Same result as doing it myself, plus the friction of opening a toolbar menu. Removed.
Extension C (reply generator). Replies were too formal for internal email and too casual for clients. Every reply needed rewriting. Removed.
Extension D (fact-checker). Returned search results, not primary sources. Not a fact-check. Removed.
A dedicated AI writing tool ($15/month). Output was only slightly better than the free tool I already use. Not worth the subscription. Cancelled.
A reply-rewriting tool. Over-apologized. Lost facts. Required heavy editing. Cancelled.
A $3/month lightweight assistant. Saved 2 minutes on one task, lost 6 minutes on another. Net negative. Cancelled.
A scheduling assistant. Got time zones wrong, ignored my preferences, and missed an incoming email. Cancelled on day three.
The pattern in all seven: they added steps instead of removing them. Every one required more correction than the tool I already had.
What I Still Don't Trust
Two tools are in a different category. I use them, but I don't trust them. They haven't earned a permanent place.
A meeting note-taker (Tool A from the background-noise test). It handled noise better than the alternatives, but it still missed three action items in one test and got a name wrong. I use it, but I re-listen to every recording. That's not trust. That's a workflow with a safety net.
A long-document summarizer. It found 11 of 12 relevant sections in a 38-page policy update. The section it missed was the one that mattered most. I use it, but I read the pages where buried details tend to live. Every time.
Both tools are useful. Neither one is trustworthy on its own. I'll test them again in a month. If they still miss the same things, I'll stop using them.

The Patterns
Here's what the month taught me about AI tools.
Simple tasks work. Complex tasks don't. The tools did best with short emails, page summaries, and straightforward meeting notes. They did worst with tasks requiring context, priority judgment, or relationship nuance.
Formatting is not evidence. The tools that looked most confident were not the most accurate. The cleanest output in the email test came from the tool that needed the most correction.
Correction time is the deciding factor. Not generation time. Not features. Not price. The tool that required the least correction was always the tool I kept. Every time.
Redaction adds real overhead. For anything involving client names or vendor data, I spent 3-5 minutes redacting before pasting. That time counts against the tool's savings.
One good tool beats five average ones. My best week used three tools total. My worst week used seven. The extra four added steps and cost me time.
What I've Changed
A month of testing changed my habits in four ways.
I redact first. Every document gets checked for names, numbers, and identifiers before it goes into any AI tool. No exceptions.
I start with a baseline. Before I test a tool, I measure how long the task takes without it. If I skip that step, I don't trust the result.
I read the output out loud. Every draft gets read aloud before it goes out. This catches tone problems the AI can't.
I write down the start and finish time. Every test. By hand. In the notebook. No estimates.
The Numbers I'd Want to See in Month Two
Here's what I'll be tracking next.
Correction time per task. I want to see if it drops as I get better at prompting.
Number of tests per week. Right now it's about 5. That's high. I want to get it to 2 or 3 and use the tools I've kept instead of testing new ones.
Time saved per week. Month one was about 2 hours. I want to see if that number is real or if it was inflated by one or two lucky tests.
Tools removed. I want to see if any of the kept tools stop earning their place. If none do, I'm not testing hard enough.
The Honest Conclusion
AI tools saved me about 2 hours a week this month. They cost me about 3 hours a week in testing, reading, correcting, and redacting. Net for the month: negative 1 hour.
That's not a failure. That's the cost of building a baseline. The tests are done. The tools are chosen. The corrections are routine. Month two should show a positive net.
But I want to be careful about that prediction. Every tool I've kept still requires review. Every output still gets read. If I stop checking, I might save 4 hours a week and get burned. The checking is not optional. It's the whole point.
Here's the rule I'm carrying into next month. It's the same rule I started with.
Test it in real life.
No comments yet.