The Demo Looked Great
Last year, a vendor demoed an AI tool for our team.
The screen was polished. The dashboard was clean. The presenter typed a prompt, and the tool produced a summary of a 40-minute meeting in about eight seconds. Everyone in the room nodded. Someone said "this could change how we work."
Then we tried it on a Tuesday.
The meeting was messier than the demo. Three people talked at once. One person had a bad microphone. Two decisions were made, but the tool labeled both as "action items." A vendor name was spelled wrong in four places. And the summary was 600 words—longer than the notes I would have written by hand.
I spent 25 minutes cleaning up the output. The demo promised eight seconds. The Tuesday reality was closer to 30 minutes, start to finish.
That experience is why this blog exists.
What This Blog Is
I'm Marcus Bell. I'm 39. I work as an operations coordinator at a midsize logistics company outside Kansas City. My day is mostly email, meeting notes, spreadsheets, scheduling, vendor follow-ups, and the small administrative tasks that quietly eat an entire workday.
I'm not an engineer. I'm not an AI researcher. I'm not a futurist. I don't have a newsletter about the future of work, and I don't want one.
What I do have is a paper notebook beside my keyboard, and a habit of writing down the start time and the finish time of every AI experiment I run at work.
That's what this blog is about. Not predictions. Tests.
What I Do All Day
Before I write about AI, I want to be clear about the work I'm testing it on.
My job is operations. That means I coordinate between teams, track shipments, follow up with vendors, handle escalations, and keep the small details from falling through the cracks. I write a lot of emails. I sit in a lot of meetings. I maintain a handful of spreadsheets that other people depend on. I write weekly reports and update a shared scheduling document that has been slowly growing for three years.
None of this is glamorous. Most of it is repetitive. And a lot of it is the kind of work where a small improvement every day adds up over a year.
That's the work I test AI on. Not toy prompts. Not abstract tasks. Actual things on my desk that need to be done by Friday.
What I'm Not Doing
Let me be specific about the boundaries.
I'm not predicting the future of work. I don't know what jobs will look like in five years. I don't know which industries will change. I don't know whether AI will replace anyone. I'm not qualified to answer those questions, and I don't think most people writing about them are either.
I'm not running a hype site. There's enough of that. I'm not going to tell you that any tool "changes everything." Most tools don't. Most tools help with one specific thing, in one specific situation, under specific conditions. That's what I want to document.
I'm not reviewing AI models in the abstract. I don't care which model "wins" a benchmark. I care whether a specific tool, used by a specific person, on a specific task, actually finished the task faster or better.
I'm not sponsored. If that changes, I'll tell you clearly. And I won't accept sponsorship from a company whose tool I've criticized in a prior test without disclosing that history.
What I Measure
Every test on this blog uses the same basic method. It's not scientific. But it's consistent, and it's honest.
The task. What was I actually trying to do? Not a fictional task. A real one. With a real deadline.
The baseline. How long would this task have taken me the normal way? I write this down before I open any AI tool. If I don't have a baseline, I don't have a test.
The tool and time. Which tool did I use? What day? What was the start time and the finish time? I record these in my notebook during the test, not after.
The output. What did the tool actually produce? I keep a copy. Not always to publish—sometimes the content is confidential—but always to compare against the final version.
The corrections. How much did I have to fix? Names, numbers, tone, structure, accuracy. This is the part most reviews skip, and it's usually the part that decides whether a tool is worth keeping.
The time saved. Not a percentage. A specific number. If I say a tool saved 12 minutes, I mean I measured it.
The verdict. Keep it, use it selectively, wait for a better version, or cancel the subscription.
The limitation. What should you not expect this tool to do? Every test ends with this. If I can't name a limitation, I haven't tested the tool hard enough.

Why the Notebook Matters
I keep a paper notebook beside my keyboard. It's nothing special—a spiral notebook I picked up at an office supply store. But it's the most important tool I use for this blog.
Here's why. When I test an AI tool, I write down:
The date and time I started
The task
The tool and version
The prompt I used
The time I finished
What I had to fix
Whether I would use it again
I write this down by hand because typing it into a spreadsheet while I'm working feels too much like the work itself. The notebook is separate. It's a record, not another app.
The notebook also keeps me honest. It's easy to remember a tool as better than it was. It's harder to argue with a start time written in ink.
What I'll Write About
The blog has five sections.
Workday Experiments. Real tests of AI on emails, meeting notes, spreadsheets, research, scheduling, and admin work. This is the core of the site.
Tool Trial Run. Structured reviews of AI tools, tested on actual tasks before a verdict is published.
Prompts That Survive Real Life. Practical prompt structures for messy inputs, vague requests, and ordinary workplace constraints. Not prompt theatre.
Outside the Office. The same testing method applied to family administration, home repairs, cooking, travel planning, and personal organization.
The Honest Debrief. Personal essays, failed experiments, ethical boundaries, and reflections on how AI changes ordinary work. This is where the blog's voice lives.
What I Won't Do
A few things you won't see here.
I won't publish a test on a tool I haven't used. If I write about it, I used it. On a real task. With a real deadline.
I won't invent productivity numbers. If I say a tool saved 15 minutes, I measured it. If I forgot to measure, I'll say so.
I won't paste confidential information into any AI tool. No client names, no vendor contracts, no coworker details, no internal project documents. If I can't test a tool without that information, I won't test it.
I won't confuse a demo with a workday. The demo is a marketing event. The workday is the test. I care about the workday.
I won't tell you that AI will replace your job. I don't know that. I don't think it's helpful to say. What I do think is helpful is telling you when a tool saved 20 minutes and when it just moved the work around.
The First Test
The first real test I'll publish is about how much time my normal workday actually loses to repetition.
Not an AI test. A baseline test. I want to know how long my normal tasks take before I start measuring what AI does to them. Otherwise every number I publish later will be a guess.
That test is the foundation. Without it, the rest of this blog is just impressions.
The Signature Question
Every post on this blog ends the same way.
Test it in real life.
That's not a slogan. It's the whole method. The demo looked good. The Tuesday test was different. Every time.
I'm not going to tell you what to think about AI. I'm going to tell you what happened when I tried it on a real task, at a real desk, with a real deadline.
If that's useful, I'm glad you're here.
No comments yet.