The Task
I had a 45-minute operations meeting on my calendar.
Five people. Two from my team, three from the warehouse side. We were working through a shipment delay that had been dragging for two weeks. The agenda was loose—no slides, no formal structure. Just a conversation that needed to end with a clear list of who was doing what by when.
I usually take notes by hand during these meetings. Afterward, I clean them up and turn them into action items. That cleanup usually takes me about 20 minutes. The whole process, from the start of the meeting to a usable summary, is roughly 35 minutes of my time.
For this test, I recorded the meeting (with consent) and gave the audio to an AI transcription tool. Then I gave the transcript to an AI summarizer. I wanted to know: would the output save me time, or would I spend more time fixing it?
The Setup
Tools tested: One AI transcription tool (paid tier) and one AI summarizer (free tier)
Date: September 24, 2026
Meeting length: 45 minutes
Participants: 5
My baseline: 35 minutes (notes + cleanup + action items)
What I wanted: A usable summary with decisions, action items, owners, and deadlines
What I measured: Time to review, number of errors, number of missing items, and final time compared to baseline
The Prompt
After transcription, I pasted the transcript into the summarizer with this prompt.
"Summarize this operations meeting. Include: (1) key decisions, (2) action items with owners, (3) unresolved questions, and (4) any dates mentioned. Keep it under 300 words. Do not add information that is not in the transcript. If something is unclear, say 'unclear' rather than guessing."
The Results
The summarizer returned output in about 12 seconds.
What it got right:
The main topic (shipment delay) and the general cause (vendor-side production issue)
The two decisions that were actually made
Three of six action items, with correct owners
The revised delivery window
The follow-up meeting date
What it got wrong:
It missed three action items entirely. One was a small task assigned to me (confirm the customer email list). Two were assigned to the warehouse team (check inventory counts and update the tracking sheet).
It misidentified one action item. A note about "checking with the vendor" was listed as an action item for the warehouse team, but it was actually a question for the vendor's account manager.
It got one name wrong. The vendor's contact was listed as "Marcus" instead of "Mark." That's my own name in the transcript—the AI confused the speaker label with the vendor name.
It listed one unresolved question as a decision. The question was "should we expedite the remaining shipment?" The transcript shows it was discussed but not decided. The AI marked it as a decision.
The Corrections
I spent 14 minutes fixing the summary.
3 minutes: adding the three missing action items
4 minutes: correcting the misidentified action item and moving it to the right owner
2 minutes: fixing the name error
3 minutes: moving the unresolved question out of decisions and into its own section
2 minutes: reading the whole thing out loud and adjusting one sentence
Total time from start to finished summary: about 27 minutes (12 seconds for AI output, 14 minutes of correction, plus setup).
My baseline was 35 minutes. The AI version took 27. That's an 8-minute saving.

What I Learned
A few things stood out.
The summary was accurate at the top level and unreliable at the detail level. It captured the main topic and the broad decisions. It missed or misread the small things—action items, owner names, and the difference between a question and a decision.
Names are still a problem. The transcript confused my name with the vendor contact's name. That's a specific kind of error that shows up in transcripts and summaries. I had to check every name against the recording.
Action items are the hardest part. The AI summarized the conversation well. But action items require knowing who said what, who owns what, and what was actually decided versus discussed. The AI didn't have that context.
The time saved was real but smaller than the demo suggested. A vendor might claim this tool saves 20 minutes per meeting. In this test, it saved 8. That's still useful—8 minutes a week adds up—but it's not a transformation.
The baseline mattered. Without my 35-minute baseline, I would have guessed the savings. My guess would have been wrong. I would have said "about 15 minutes" and been off by nearly double.
What Still Needed My Attention
Even after the corrections, I still did three things manually.
I checked every name against the recording. Two minutes of listening at specific timestamps. The AI cannot be trusted with names.
I confirmed every action item against the transcript. I re-read the relevant sections to make sure each owner was correct.
I made the final call on the unresolved question. The AI wanted to list it as a decision. I moved it to unresolved. That's a judgment call, not a generation task.
The Verdict
Kept, with a specific use case. I'll keep using the transcription and summary tools, but only for the first draft. The output always needs a review pass for names, action items, and decision-versus-question distinctions.
The 8-minute saving is real. On a weekly meeting, that's about 35 minutes a month. Not huge, but not nothing. And the summary quality is better than what I'd produce if I were rushed and trying to write notes while listening.
Baseline note: My baseline for this task was 35 minutes. The AI version came in at 27. That's about a 23% reduction. I'll track this again over the next month to see if the pattern holds.
The Limitation
One meeting is not enough to judge this workflow.
This was a loose, conversational meeting with five people. A more structured meeting with an agenda might produce a more accurate summary. A less structured meeting might produce a worse one. I'll test both before drawing a stronger conclusion.
Also, the transcript tool was paid. The summarizer was free. Paid summarizers might handle names and action items better. I haven't tested that yet.
And I used a recording. If I were taking notes in real time and feeding those notes to the AI, the result would be different. That's a different test.
Test it in real life.
No comments yet.