> ## Content Index
> Fetch the complete content index at: https://www.thedailyconstraint.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Your Automation “Feels” Like It's Working. That's Not a KPI.
- URL: https://www.thedailyconstraint.com/your-automation-feels-like-its-working-thats-not-a-kpi/
- Published: 2026-09-03T11:06:26.000Z
- Updated: 2026-09-04T02:40:46.000Z
- Author: Andrea Bullock

Someone in the status meeting says the new AI tool is working great.

You ask what “great” means. Numbers, a rate, anything.

You get a shrug, and a story about a really smooth Tuesday a few weeks back.

*No one had a number. They had a feeling.*

A feeling can't be tracked. It can't be compared week over week. It definitely can't tell you the exact moment things started quietly getting worse.

It just tells you how the room felt in that meeting, about a Tuesday nobody measured.

Standard work tells the tool what its job is. It doesn't tell you whether the tool is doing that job well. LINK: [Your AI Needs Standard Work Too](https://www.thedailyconstraint.com/your-ai-needs-standard-work-too/).

**Why “It Feels Like It's Working” Isn't Data**

A feeling has no denominator. Nobody can say “it feels like it's working” 94 percent of the time, or point to which 6 percent it isn't.

It also has no memory. A rough Thursday two months ago and a rough Thursday last week both just blend into “mostly fine.”

Real measurement catches drift early. A feeling catches it only once it's bad enough that someone finally complains out loud.

**The Real Metrics: Completion Rate, Accuracy, Intervention Rate, Rework**

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/08/B1-6.png)

None of these require a data science team. They require someone willing to count.

● Completion rate: the share of tasks the tool finishes end to end, with a real [definition of done](https://www.scrum.org/resources/what-definition-done?ref=thedailyconstraint.com) behind it, not just what it reports as finished.

● Accuracy: of the tasks it completes, how many are correct, checked against a real outcome, not just against its own confidence.

● Intervention rate: how often a human had to step in, at any checkpoint, for any reason.

● [Rework](https://asq.org/quality-resources/cost-of-quality?ref=thedailyconstraint.com): how much of what the tool produced had to be redone by a person afterward, and how long that took.

Track all four, even loosely, and you already know more about whether this tool is working than most rollout status meetings ever ask.

**Goal Success Isn't the Same as Step-by-Step Accuracy**

It's worth separating two things people tend to blur together. Whether the overall task got done, and [whether every step along the way was correct](https://arize.com/guides/ai-agent-handbook/agent-evaluation-metrics/?ref=thedailyconstraint.com).

A tool can hit its overall goal while getting the middle badly wrong, the same way an approval workflow can technically finish while quietly repeating one bad assumption nine times over.

Measuring only the finish line hides exactly that kind of problem. Both numbers matter, and they're not the same number.

**This Is the Same Work an Analyst Used to Do for You**

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/08/B2-6.png)

[Pulling these numbers together, cleanly, on a regular basis, used to require someone with real analyst skills](https://www.thedailyconstraint.com/what-can-an-experienced-operations-manager-do-with-ai-that-used-to-require-an-analyst/). An AI tool can help build this scorecard the same way it can help build the automation itself.

Turning the tool on it doesn't feel as satisfying as the demo did. It's the only part of this whole category that tells you whether the demo was right.

**The Prompt: Build a Real Scorecard for Your AI Tool**

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/08/B3-4.png)

Someone in the status meeting says the new AI tool is working great.

You ask what “great” means. Numbers, a rate, anything.

You get a shrug, and a story about a really smooth Tuesday a few weeks back.

*No one had a number. They had a feeling.*

A feeling can't be tracked. It can't be compared week over week. It definitely can't tell you the exact moment things started quietly getting worse.

It just tells you how the room felt in that meeting, about a Tuesday nobody measured.

Standard work tells the tool what its job is. It doesn't tell you whether the tool is doing that job well. \[LINK: Your AI Needs Standard Work Too\]

**Why “It Feels Like It's Working” Isn't Data**

A feeling has no denominator. Nobody can say “it feels like it's working” 94 percent of the time, or point to which 6 percent it isn't.

It also has no memory. A rough Thursday two months ago and a rough Thursday last week both just blend into “mostly fine.”

Real measurement catches drift early. A feeling catches it only once it's bad enough that someone finally complains out loud.

**The Real Metrics: Completion Rate, Accuracy, Intervention Rate, Rework**

None of these require a data science team. They require someone willing to count.

● Completion rate: the share of tasks the tool finishes end to end, with a real [definition of done](https://www.scrum.org/resources/what-definition-done?ref=thedailyconstraint.com) behind it, not just what it reports as finished.

● Accuracy: of the tasks it completes, how many are correct, checked against a real outcome, not just against its own confidence.

● Intervention rate: how often a human had to step in, at any checkpoint, for any reason.

● [Rework](https://asq.org/quality-resources/cost-of-quality?ref=thedailyconstraint.com): how much of what the tool produced had to be redone by a person afterward, and how long that took.

Track all four, even loosely, and you already know more about whether this tool is working than most rollout status meetings ever ask.

**Goal Success Isn't the Same as Step-by-Step Accuracy**

It's worth separating two things people tend to blur together. Whether the overall task got done, and [whether every step along the way was correct](https://arize.com/guides/ai-agent-handbook/agent-evaluation-metrics/?ref=thedailyconstraint.com).

A tool can hit its overall goal while getting the middle badly wrong, the same way an approval workflow can technically finish while quietly repeating one bad assumption nine times over.

Measuring only the finish line hides exactly that kind of problem. Both numbers matter, and they're not the same number.

**This Is the Same Work an Analyst Used to Do for You**

[Pulling these numbers together, cleanly, on a regular basis, used to require someone with real analyst skills](https://www.thedailyconstraint.com/what-can-an-experienced-operations-manager-do-with-ai-that-used-to-require-an-analyst/). An AI tool can help build this scorecard the same way it can help build the automation itself.

Turning the tool on it doesn't feel as satisfying as the demo did. It's the only part of this whole category that tells you whether the demo was right.

**The Prompt: Build a Real Scorecard for Your AI Tool**

Give an AI tool, Claude, ChatGPT with data analysis enabled, or similar, whatever logs or output records you already have from the tool you're evaluating. Ask it to help you build a real scorecard instead of relying on impressions. Then adapt this prompt to your situation.

**1\. Define Completion Rate for This Specific Task**

● State exactly what counts as a completed task, using the definition of done you already set for this process.

● Have it calculate the rate from whatever records you can provide.

**2\. Define Accuracy Separately From Completion**

● Ask it to help you sample completed tasks and check them against the actual correct outcome.

● Example instruction: “Help me design a sampling method to check accuracy without reviewing every single case.”

**3\. Track the Intervention Rate**

● Count how often a human had to step in at any checkpoint, and why.

● A rising intervention rate is often the earliest sign something's drifting, well before completion rate ever drops.

**4\. Track Rework, Not Just Errors**

● Measure how much of the output required real correction afterward, and roughly how long that correction took.

● A small error rate can still hide a large rework burden if each fix is expensive.

**5\. Set a Review Cadence, Not a One-Time Check**

● Ask it to suggest how often these four numbers should get reviewed, based on how much volume this process handles.

● A scorecard nobody looks at again is just a nicer-looking version of the same shrug from the status meeting.

Five steps, and what you get is an answer that holds up when someone asks a hard question about it. “It feels fine” never does.

**What AI Can't Score for You**

AI can calculate all four numbers accurately, all day, without getting tired of counting.

It can't decide what threshold matters for your operation. A 90 percent completion rate might be excellent for one process and alarming for another.

It also can't feel the difference between a number that's technically fine and one that's about to become a real problem the moment volume spikes.

The scorecard tells you what happened. Deciding what to do about it is still entirely yours.

**Want the Short Version?**

If you'd rather see this argument play out rather than read through the numbers again, watch this video. It explains what can happen when an AI tool "feels" like it's working, but no one is measuring it. 

Sometimes, the first sign your automation is failing isn't an error message. It's the quiet workaround no one thought to count.

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/09/image.png)

**Next Up in This Category**

A good scorecard answers whether the tool is working today.

It doesn't answer what happens to that tool three months from now, once the launch excitement has worn off and nobody's checking anymore. \[LINK: AI Automation Isn't “Set It and Forget It”\]

Subscribe to The Daily Constraint if you want it the day it goes live.

Do you have numbers behind your AI tool's performance, or a feeling about a good Tuesday? Tell me which one you're working with.

Give an AI tool, Claude, ChatGPT with data analysis enabled, or similar, whatever logs or output records you already have from the tool you're evaluating. Ask it to help you build a real scorecard instead of relying on impressions. Then adapt this prompt to your situation.

**1\. Define Completion Rate for This Specific Task**

● State exactly what counts as a completed task, using the definition of done you already set for this process.

● Have it calculate the rate from whatever records you can provide.

**2\. Define Accuracy Separately From Completion**

● Ask it to help you sample completed tasks and check them against the actual correct outcome.

● Example instruction: “Help me design a sampling method to check accuracy without reviewing every single case.”

**3\. Track the Intervention Rate**

● Count how often a human had to step in at any checkpoint, and why.

● A rising intervention rate is often the earliest sign something's drifting, well before completion rate ever drops.

**4\. Track Rework, Not Just Errors**

● Measure how much of the output required real correction afterward, and roughly how long that correction took.

● A small error rate can still hide a large rework burden if each fix is expensive.

**5\. Set a Review Cadence, Not a One-Time Check**

● Ask it to suggest how often these four numbers should get reviewed, based on how much volume this process handles.

● A scorecard nobody looks at again is just a nicer-looking version of the same shrug from the status meeting.

Five steps, and what you get is an answer that holds up when someone asks a hard question about it. “It feels fine” never does.

**What AI Can't Score for You**

AI can calculate all four numbers accurately, all day, without getting tired of counting.

It can't decide what threshold matters for your operation. A 90 percent completion rate might be excellent for one process and alarming for another.

It also can't feel the difference between a number that's technically fine and one that's about to become a real problem the moment volume spikes.

The scorecard tells you what happened. Deciding what to do about it is still entirely yours.

**Next Up in This Category**

A good scorecard answers whether the tool is working today.

It doesn't answer what happens to that tool three months from now, once the launch excitement has worn off and nobody's checking anymore. \[LINK: AI Automation Isn't “Set It and Forget It”\]

Subscribe to The Daily Constraint if you want it the day it goes live.

Do you have numbers behind your AI tool's performance, or a feeling about a good Tuesday? Tell me which one you're working with.