> ## Content Index
> Fetch the complete content index at: https://www.thedailyconstraint.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Your Pilot Isn't a Test: How AI Builds One That Actually Proves Something
- URL: https://www.thedailyconstraint.com/your-pilot-isnt-a-test-how-ai-builds-one-that-actually-proves-something/
- Published: 2026-08-26T15:18:30.000Z
- Updated: 2026-08-28T13:12:36.000Z
- Author: Andrea Bullock

So, the fix looks safe to roll out.

The math checks out. The downstream station has room.

Time to just try it, right?

Please don't.

A model that says a fix is safe is not the same thing as proof that it works.

You still don't know whether the actual signoff change does what the spreadsheet says it should. Not with real operators. Not on a real shift.

That's what a pilot is for. *“Just try it” is not a pilot.*

Last time, the model showed you where the Line 4 fix would push the strain once it sped things up. Dock capacity, sitting closer to its limit than anyone realized. If you missed how we got there, catch up here. [You Didn't Fix the Bottleneck: How AI Predicts Where It Moves Next](https://www.thedailyconstraint.com/you-didnt-fix-the-bottleneck-how-ai-predicts-where-it-moves-next/).

**Why “Just Try It” Isn't a Test**

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/08/Blog-1-2.png)

You already know how this usually goes. Someone proposes a fix. The room nods.

Somebody runs it for a week. Three people say it seems better. Nobody wrote down what “better” meant before they started.

Two months later, you're arguing about whether it worked. Using nothing but memory, vibes, and whoever complained loudest that week.

That's not a test. That's a guess with a start date.

A real pilot has structure. The structure is what separates “I think this worked” from “I can prove this worked.” That's the same standard your hypothesis had to meet, a few pieces back in this series.

**The Nine Things a Real Pilot Needs**

![](https://storage.ghost.io/c/6a/5a/6a5a81c1-c40d-4db0-af99-e0cedc71c0c2/content/images/2026/08/blog-2.png)

This looks like a lot up front. It isn't, once AI is doing the heavy lifting. Here's the full list, in the order you'd build it:

**1\. Baseline**

● What's happening right now, before you touch anything.

● Skip this and you'll be arguing forever about whether the fix worked, with nothing to argue from.

**2\. Intervention**

● The specific fix, at a specific station, for a specific window of time.

● Not “we're trying some AI stuff.” The actual sign-off step change on Line 4, starting Monday, running for three weeks.

**3\. Control Variables**

● Everything you're deliberately holding steady.

● This is how you know the fix is what changed, and not something else that happened to shift at the same time.

**4\. Confounders**

● The sneaky stuff that could explain your results, even if the fix did nothing at all.

A [confounding variable](https://www.simplypsychology.org/pilot-studies.html?ref=thedailyconstraint.com) is any outside factor that changes alongside your fix. It offers a competing explanation for whatever you see.

Catching these before you start is most of what separates a real pilot from an expensive guess.

Here's one that gets almost everyone, at least once. Run a pilot on one line. Every operator on it knows they're being watched.

That's the [Hawthorne effect](https://www.scribbr.com/research-bias/hawthorne-effect/?ref=thedailyconstraint.com): the well-documented tendency for people to work differently, the moment they know someone's checking.

Congratulations. You may have just built the least controlled experiment in the history of science.

Your “successful pilot” might just be four operators having their sharpest three weeks of the year, because the plant manager keeps walking by.

**5\. Leading Indicators**

● An early signal that tells you within days, not months, whether this is heading somewhere good.

**6\. Lagging Indicators**

● The slower number that shows up later, and confirms the leading indicator wasn't a fluke.

[The distinction matters more than it sounds](https://www.advancedtech.com/blog/leading-vs-lagging-kpis/?ref=thedailyconstraint.com).

Lean on lagging indicators alone. You won't know your pilot is failing, until it's already over.

A leading indicator is what lets you catch trouble in week one instead of week three.

**Applied to Line 4**

Stripped of the theory, here's roughly what those nine pieces look like for the actual changeover fix you're testing:

● Baseline: changeover time and failure rate on Line 4 over the last two weeks, before anyone touches the signoff step.

● Intervention: mandatory sign-off confirmation, enforced on every changeover, for three weeks.

● Leading indicator: signoffs completed and logged, checked daily. If that number isn't near 100 percent by day three, nothing downstream will matter.

● Lagging indicator: changeover-related downtime, measured weekly across the full three-week run.

● Stop condition: downtime on the receiving station spikes past a level you set in advance. That's exactly where the constraint migration model said the strain would land.

● Success threshold: downtime drops by a set amount, without a matching increase anywhere downstream. Decided and written down before the pilot starts.

None of that required a data scientist. It required someone willing to write the numbers down before starting. That's the entire discipline in one sentence.

**7\. Stop Conditions**

● The specific line you draw in advance, written down, for when you'll kill the pilot early.

● Decide this before you start. Otherwise a bad test just quietly runs forever, because nobody wants to be the one who calls it.

**8\. Sample Requirements**

● How many changeovers, shifts, or cycles you need before you have enough data to trust the result, good or bad.

**9\. Success Threshold**

● The number that means “roll it out everywhere,” decided before you see any results.

● Set this after the data comes in, and you'll find yourself explaining, with a straight face, why a two percent improvement was the goal all along.

**The Prompt: Design Your Pilot**

Give an AI tool the specifics of your countermeasure, and the process it's changing: Claude, ChatGPT with data analysis enabled, or similar. Ask it to help you build all nine pieces before you touch anything on the floor. Then adapt this prompt to your situation.

**1\. Describe the Countermeasure Precisely**

● State exactly what's changing, where, and for how long.

● Vague inputs produce a vague pilot design.

**2\. Ask It to Draft the Baseline Measurement Plan First**

● Before anything else, have it specify exactly what to measure, and for how long, to establish a credible “before” picture.

● Example instruction: “Before we design the pilot itself, tell me what baseline data I need and how many shifts I should collect it over to make it credible.”

**3\. Have It Separate Leading From Lagging Indicators**

● Ask for at least one of each.

● Ask it to explain why each one qualifies as leading or lagging, for this specific process.

**4\. Force It to Name the Confounders, Including the Awkward Ones**

● Explicitly ask what could make this pilot look successful even if the fix does nothing.

● Include observation effects, seasonal patterns, or anything else specific to your operation.

**5\. Make It Commit to a Stop Condition and a Success Threshold, in Writing**

● Both numbers, before the pilot starts.

● Ask it to flag if either one is vague enough to argue with later.

**6\. Ask for the Sample Size, and Ask It to Show Its Work**

● Have it explain, in plain terms, why that many cycles or shifts are enough to trust the result.

● Not just hand you a number.

Run all six, and you've got a pilot design that can survive someone asking hard questions about it.

**What AI Can't Judge for You**

AI can build a structurally sound pilot design in minutes.

It cannot decide how much risk your operation can tolerate while that pilot runs. It cannot judge whether three weeks is politically survivable, if the fix makes things temporarily worse before it gets better.

It cannot read the room when your team is exhausted from the last three changes. They need this one to go smoothly, even if smoothly means slower.

And it cannot walk the floor on day one. It won't notice the operators quietly working around the new sign-off step instead of following it. That alone would wreck the whole pilot before it collects a single useful data point.

The design is the easy part. Running it on a real floor, with real people, while real orders still need to ship, is still entirely yours.

**Next Up in the Series**

Say the pilot clears its threshold. You're ready to roll the fix out everywhere.

Before you do, it's worth one more gut check. Everything that could go wrong, once this becomes standard practice instead of a pilot.

Next up, we use AI to run a pre-mortem on the full rollout. Attacking your own plan from every angle before the floor finds those angles for you. LINK: [Nobody Raised Their Hand: How AI Finds the Risks Your Rollout Meeting Missed](https://www.thedailyconstraint.com/how-can-operations-managers-use-ai-to-run-a-pre-mortem-before-changing-a-process/).

Subscribe to The Daily Constraint. It'll land in your inbox the day it goes live.

Design a pilot this week? Comment below. Tell me what your stop condition was. Not whether the pilot worked. Whether you'd already decided what would make you kill it.