The Problem with Testing Without a Baseline
Before I test any AI tool, I need to know what the normal version looks like.
Not the version I remember. Not the version I estimate. The version I measured.
I learned this the hard way last year. I tested an AI email tool and wrote that it saved me "about 20 minutes a day." Two weeks later, I tried to verify that number and realized I had no idea where it came from. I had never measured how long my email took before the tool. I had guessed.
That guess turned out to be wrong. My normal email time was closer to 35 minutes a day, and the tool saved me closer to 12.
So before I test anything for this blog, I want a real baseline. Here's how I built one.
The Setup
What I tracked: My normal workday, at my actual job.
When: One full work week (five days).
How: A paper notebook beside my keyboard. Every time I started a task, I wrote the time. Every time I finished, I wrote the time. I also wrote a one-word task label.
What I didn't do: I didn't use any AI tools during this week. I didn't change my normal routine. I wanted the honest version of my week.
Job context: Operations coordinator at a midsize logistics company outside Kansas City.
Work hours: Roughly 8:15 AM to 5:00 PM, with a 30-minute lunch.
The Categories
I grouped every tracked task into one of six categories.
1. Email. Reading, drafting, replying, following up, organizing.
2. Meeting notes. Writing notes during meetings, cleaning them up afterward, sending summaries.
3. Spreadsheets. Updating tracking sheets, cleaning data, building lists.
4. Reports. Weekly reports, monthly reports, ad-hoc reports for managers.
5. Scheduling. Coordinating calendars, confirming times, rescheduling.
6. Vendor follow-ups. Chasing quotes, confirming deliveries, resolving issues.
Those six categories cover about 80% of my workday. The remaining 20% is meetings themselves, which I didn't track for this baseline because AI doesn't change that time yet.
The Numbers
Here's what one week of tracking looked like.
Total tracked time: 32 hours 45 minutes across five days.
Email: 9 hours 20 minutes.
Meeting notes: 5 hours 10 minutes.
Spreadsheets: 4 hours 45 minutes.
Reports: 4 hours 30 minutes.
Scheduling: 3 hours 20 minutes.
Vendor follow-ups: 3 hours 15 minutes.
Other (untracked): 2 hours 25 minutes.
The email number surprised me the most. I thought email was maybe 6 hours a week. It was closer to 9.5. That's almost 20% of my tracked time on a single category.
The meeting notes number also surprised me. I thought of it as a small task. But it was over an hour a day, spread across writing during meetings and cleaning up afterward.
What Counts as "Repetition"
The question I wanted to answer wasn't "how long do tasks take." It was "how much of that time is repetition."
I defined repetition as work that fits all three of these conditions:
The task follows the same general structure every time.
The output varies, but the process doesn't.
I could describe the steps to a new employee in under five minutes.
Using that definition, here's how much of each category was repetitive:
Category | Total time | Repetitive time | Repetitive % |
|---|---|---|---|
9h 20m | 6h 10m | 66% | |
Meeting notes | 5h 10m | 4h 05m | 79% |
Spreadsheets | 4h 45m | 3h 20m | 70% |
Reports | 4h 30m | 3h 15m | 72% |
Scheduling | 3h 20m | 2h 40m | 80% |
Vendor follow-ups | 3h 15m | 2h 30m | 77% |
Total | 32h 45m | 22h 00m | 67% |
That's the number I was looking for. About two-thirds of my tracked workday fits the definition of repetition.
Not all of that can be delegated to AI. Some of it needs judgment. Some of it needs relationship context. Some of it is repetitive but too small to automate usefully.
But it gives me a ceiling. There are about 22 hours a week of repetitive work in my schedule. That's the maximum any AI tool could theoretically address, and the number I'll measure against in every future test.

What I Learned
A few things stood out.
My guesses were wrong. Every time. Email was higher. Meeting notes were higher. Scheduling was lower. If I had published those guesses as "estimates," they would have been misleading.
Repetition is not the same as low value. A lot of my repetitive work matters. Vendor follow-ups, for example, are repetitive but essential. Saving time there doesn't mean the work is unimportant.
The highest-repetition category was the smallest. Scheduling was 80% repetitive but only 3 hours a week. The savings ceiling is small.
The largest opportunity was in the biggest category. Email was 9.5 hours and 66% repetitive. That's the category most likely to benefit from a tool, because the total time is high.
Tracking itself cost time. Recording start and finish times added maybe 15 minutes a day. That's a real cost. I'm not going to track every task forever. I'll do it during test weeks only.
The Baseline I'll Use
Here's what I'm taking forward from this test.
Baseline per category:
Email: 9h 20m per week
Meeting notes: 5h 10m per week
Spreadsheets: 4h 45m per week
Reports: 4h 30m per week
Scheduling: 3h 20m per week
Vendor follow-ups: 3h 15m per week
Repetitive ceiling per category:
Email: 6h 10m per week
Meeting notes: 4h 05m per week
Spreadsheets: 3h 20m per week
Reports: 3h 15m per week
Scheduling: 2h 40m per week
Vendor follow-ups: 2h 30m per week
Any AI test I publish from here will be measured against these numbers. If an email tool saves 12 minutes a day, that's 1 hour per week against a 6-hour ceiling. That's a real number, and I can compare it to the next tool.
The Limitation
This baseline is mine. It's not universal.
My job is operations at a midsize logistics company. Someone in sales, or healthcare, or education, or a small agency would have a different baseline. The specific numbers here are not generalizable. The method is.
I also only tracked one week. One week isn't enough to smooth out unusual weeks. A monthly baseline would be better. I may rebuild this every few months as a check.
And I didn't track task quality. This baseline measures time, not output. If I save time by producing worse work, I haven't saved anything. Every future test will need to check both.
What Comes Next
The first real tool test on this blog will use this baseline.
I'll test one tool on one category—likely email first, since it's the biggest. I'll measure the starting time, the corrections, the output quality, and the finish time. I'll compare it to 9h 20m.
If the tool saves 20 minutes a week, that's about 3% of my total tracked time. If it saves 2 hours, that's closer to 10%. Either number is fine. What matters is that it's measured, not guessed.
Test it in real life.
No comments yet.