The Task
Last November, I had a vendor question.
One of our carriers had changed their fuel surcharge policy. The old policy was a flat percentage. The new policy was tiered by distance. I needed to know the new tiers so I could update our shipping cost spreadsheet.
I didn't want to call the carrier. Their hold times are legendary. I didn't want to dig through the contract PDF either—it's 47 pages and the surcharge section is buried in an appendix. So I asked an AI tool.
That was my first mistake.
The Setup
Tool: A general-purpose AI assistant (free tier)
Date: November 14, 2025
Task: Find the updated fuel surcharge tiers for a specific carrier
Input I provided: A short question: "What are the current fuel surcharge tiers for [carrier name]?"
What I expected: A list of tiers with distance ranges and percentages
What I did next: I copied the answer into our shipping spreadsheet and updated the formulas
I did not check the carrier's website. I did not open the contract PDF. I did not call the carrier. I trusted the answer because it looked specific and it was formatted neatly.
The Answer
The AI gave me a clean table. Four tiers. Distance ranges. Percentages. It looked like this:
Distance | Surcharge |
|---|---|
0–500 miles | 6.5% |
501–1,000 miles | 7.2% |
1,001–1,500 miles | 8.0% |
1,501+ miles | 9.1% |
It was confident. It was structured. It even added a note: "These rates are current as of 2024."
I didn't question it. I pasted it into the spreadsheet. I updated the formulas. I sent the updated file to my manager.
The Problem
Three weeks later, a customer invoice came back with a dispute.
The customer had been charged a 6.5% fuel surcharge on a 1,400-mile shipment. According to my spreadsheet, that distance fell into the 8.0% tier. According to the customer, it should have been 5.2%.
I pulled up the carrier's actual policy. The real tiers were different. The real surcharge for 1,001–1,500 miles was 5.2%, not 8.0%. The AI had invented the entire table.
Not the format. Not the structure. The numbers. All four tiers were wrong. The AI had produced a plausible-looking table that had no relationship to the carrier's actual policy.

The Cost
The invoice dispute was small. The difference on one shipment was about $40. I fixed it in a few minutes.
But the ripple effects were bigger.
Trust. My manager asked how the error happened. I had to explain that I'd used an AI tool and hadn't checked it. That's not a comfortable conversation.
Time. I spent about two hours re-checking every formula in the spreadsheet. If the fuel surcharge tiers were wrong, what else was wrong? I went through the whole file line by line. Nothing else was broken, but I couldn't be sure until I checked.
Process. I had to rebuild my own credibility with the manager. I promised I'd change my process. That promise is why I started keeping a paper notebook and testing AI answers before using them.
Confidence. For weeks afterward, I second-guessed every AI answer I got, even simple ones. That's not a bad thing in the long run—it made me more careful—but it was exhausting in the short term.
What I Learned
Here's what that experience taught me.
A confident answer is not a correct answer. The AI didn't say "I'm not sure." It didn't say "please verify this." It gave me a clean table and a date. It sounded authoritative. It was wrong.
Formatting is not evidence. The table looked real because it was formatted like a real table. But the formatting was just a shape. It didn't mean the numbers were true.
"Current as of 2024" is a red flag. The AI added a date to make the answer seem verified. But the carrier had changed its policy in 2025. The AI was using outdated—or invented—information.
The most dangerous errors hide in specifics. If the AI had said "fuel surcharges vary," I would have looked them up. But it gave me specific tiers with specific percentages. That specificity is what made me trust it.
I was tired. It was a Thursday. I had three other tasks on my plate. I took the shortcut because I was busy. The shortcut cost me more time than the long way would have.
The fix wasn't "stop using AI." The fix was "check the answer before using it." I still use AI every day. I just don't paste the output into a spreadsheet without verifying it first.
What I Do Now
After that mistake, I built a simple three-step check.
Step 1: Find the source. If the AI cites a source, I open it. If the AI doesn't cite a source, I find one myself. No source, no trust.
Step 2: Compare the numbers. If the AI gives me specific figures, I compare them to the primary source. Even one match is not enough—I check at least two.
Step 3: Check the date. If the AI includes a date, I verify it. If the source is older than six months, I look for a newer version.
That's it. Three steps. It takes about five minutes for a table like the one I got. It would have caught this error immediately.
I also write down the tool, the date, the question, and the answer in my paper notebook. That way, if I ever need to explain how I arrived at a number, I can. And if I find an error later, I can trace it back.
The Limitation
This process doesn't eliminate all errors.
It catches the obvious ones—wrong numbers, missing sources, outdated information. It doesn't catch errors in interpretation. If I ask the AI to summarize a policy and it summarizes it wrong, the source might still support the summary on the surface. I'd need to read the whole policy to catch that.
It also doesn't fix the trust problem. Even with the three-step check, I now trust AI answers less than I did before. That's not a bad thing, but it's a real change. I move a little slower. I second-guess a little more. That's the cost of having been burned.
The process also takes time. Five minutes per table adds up over a week. But five minutes is less than two hours of fixing a mistake and rebuilding trust. The math works.
The Bottom Line
The first time an AI tool gave me confidently wrong information, I didn't catch it.
The second time, I did. Because I checked.
That's the whole story. No dramatic conclusion. No promise that AI is useless. Just a reminder that a confident answer is not the same as a correct one. And that checking takes less time than fixing.
Test it in real life.
No comments yet.