DeskProof AI
Tool Trial Run 2026-09-23 20:43 7 reads

I Tested Three AI Meeting Note-Takers With Background Noise

I Tested Three AI Meeting Note-Takers With Background Noise

I tested three AI meeting note-takers in a real office with real background noise. One handled it well. Two did not. Here's what I measured.

The Task

I work in an open office.

That means the "quiet" meeting rooms are not always quiet. There's a coffee machine that grinds beans every few minutes. Someone on the sales team is always on a call. The HVAC system hums in the background all day. And on Tuesdays, the entire floor smells like someone's reheated fish.

When I started testing AI meeting note-takers, I realized my test conditions were unrealistic. I was testing in a quiet room with a clear microphone and no interruptions. Real meetings don't work like that.

So I ran a different test. Same three tools. Same meeting type. Same five participants. But this time, in a room with real background noise.

The Setup

Tools tested: Three AI meeting note-takers (free trials or free tiers)

Date: September 29, 2026

Meeting length: 40 minutes

Participants: 5 (2 from my team, 3 from another department)

Test location: A mid-size conference room with a window facing a parking lot and a coffee machine outside the door

Background noise present:

  • Coffee grinder running twice during the meeting

  • HVAC hum

  • Occasional hallway conversation

  • A cell phone vibrating on the table

  • One participant with a slight echo on their microphone

My baseline: 30 minutes (handwritten notes plus cleanup)

What I measured: Transcription accuracy, speaker identification, action-item accuracy, and total correction time

The Three Tools

I'll call them Tool A, Tool B, and Tool C.

Tool A is a dedicated meeting note-taker that connects to calendar invites and records directly. Paid tier.

Tool B is a general-purpose transcription service that I used as a standalone recorder. Free tier.

Tool C is a meeting assistant built into a video conferencing platform. Free with the platform.

The Results

Here's what each tool produced.

Tool A

Tool A handled the background noise best. It correctly identified four of five speakers. It captured the two coffee grinder moments without misattributing anything. When the HVAC hum was loudest, it missed two sentences, but flagged them as "unclear" instead of guessing.

But Tool A got three action items wrong. It attributed one to the wrong owner. It missed one deadline entirely. And it labeled a general discussion as an action item.

Correction time: 11 minutes.

Tool B

Tool B struggled with the background noise. It misattributed two sentences to the wrong speaker during the coffee grinder moment. It also transcribed the HVAC hum as a brief burst of unclear text in two places, which I had to delete.

The action items were mixed. It caught four of six correctly. The other two were either missing or vague.

Correction time: 16 minutes.

Tool C

Tool C was the weakest on background noise. It lost about 90 seconds of the meeting during the two coffee grinder moments. It merged two speakers into one for most of the second half. And it missed two of six action items.

It did one thing better than the others. Its summary was the most readable. The structure was clean and the flow was smooth. But it wasn't accurate enough to use without a full re-listen.

Correction time: 22 minutes.

The Comparison

Tool

Transcription accuracy

Speaker ID

Action items correct

Correction time

Tool A

High

4/5 speakers

3/6

11 min

Tool B

Medium

3/5 speakers

4/6

16 min

Tool C

Low

2/5 speakers

2/6

22 min

The pattern was consistent. The tool that handled background noise best also produced the most useful output. Noise tolerance mattered more than any other feature.

Hands marking a three-column AI tool comparison table with checkmarks and X marks in a notebook beside a laptop.

What I Learned

A few things stood out.

Background noise is the real test. A tool that works in a quiet room might not work in an open office. The tools that can't distinguish between speech and ambient noise are not useful for real work.

The "unclear" flag is a feature, not a bug. Tool A flagged the sentences it couldn't hear clearly. That was more useful than guessing. I knew exactly where to listen to the recording. Tool B and Tool C guessed, and I had to re-listen to the whole thing.

Speaker identification matters more than I expected. When two people are talking, knowing who said what is essential for action items. Tools that merged speakers made the action items unreliable.

Correction time is the deciding factor. All three tools "worked" in the sense that they produced a transcript. But the time it took to fix the output varied by a factor of two. The tool that took 11 minutes to fix saved me real time. The tool that took 22 minutes didn't save anything.

My baseline of 30 minutes was still relevant. Tool A saved me 19 minutes. Tool B saved 14. Tool C saved 8. Only one of those numbers is worth repeating.

What Still Needed My Attention

Every tool required the same manual checks.

Re-listening to noisy sections. I re-listened to the two coffee grinder moments and confirmed what was actually said. Every tool required this.

Verifying action items. I read every action item and cross-checked against my memory of the conversation. This took the most time.

Confirming names. None of the tools got every name right. I corrected three names across the three tests.

Final read. I read every summary out loud before using it. That's a check no tool does.

The Verdict

Tool A — Kept. Best noise tolerance. Best speaker identification. Least correction time. I'll keep using it for meetings in noisy rooms. The 11-minute correction is still not ideal, but it's the best of the three.

Tool B — Kept, conditionally. I'll use it for meetings in quiet rooms only. In noisy environments, it's not reliable enough to save time.

Tool C — Not kept for this use case. The output was readable, but the accuracy wasn't good enough to justify using it over taking notes by hand. I might use it for pure video calls where everyone is on their own microphone.

Baseline note: My baseline was 30 minutes. Tool A came in at 11 minutes of correction plus about 4 minutes of setup. That's a net saving of about 15 minutes. Tool B saved about 10. Tool C saved about 4. Only Tool A is worth the subscription in a noisy office.

The Limitation

This test used one meeting in one room on one day.

A quieter room might change the results. A meeting with fewer participants might change the results. A meeting with all participants on individual microphones would definitely change the results.

I also didn't test paid tiers on every tool. Tool B and Tool C were free versions. Paid versions might handle noise better. I'll test that separately before making a stronger recommendation.

And one more limitation: I didn't test a recording tool with an external microphone. A lapel mic or a conference room mic might produce better input than a laptop mic. That's worth testing next.

Test it in real life.

Last updated — 2026-09-23 20:44
Comments [ 0 ]

No comments yet.

Leave a comment