The 5 Whys Are Broken: How AI Helps You Track Multiple Root Causes

Share
The 5 Whys Are Broken: How AI Helps You Track Multiple Root Causes

Line 4 goes down. Someone asks why, five times in a row.

By the fifth why, the room has an answer. Everyone nods. The meeting ends.

The problem comes back in six weeks.

Wearing a slightly different hat.

Last time, we used AI to pull a short list of leads on Line 4's mystery downtime straight out of the data. Ranked. Sourced. If you missed it, catch up here. LINK: How Can Operations Managers Use AI to Find Relationships Their KPI Dashboard Doesn't Show?

This is the next step. Turn the strongest lead into something you can test. Not something you just believe.

A lead isn't a root cause. It's a suspect.

The tool most operations managers reach for next is the 5 Whys. It was never built to interrogate more than one suspect at a time.

Why the 5 Whys Keeps Failing You

It fails for two reasons. Only one of them is about the method itself:

●       Structural: The 5 Whys was built to follow one chain of cause and effect to one answer. Most operational failures aren't one answer. They're several conditions that only cause a failure when they show up together, the exact kind of pattern the last piece was built to catch.

●       Human: Once a room lands on an answer that feels right, everyone stops looking. Contradicting evidence gets waved off. Alternative explanations never get named out loud, let alone tested.

That's confirmation bias. It's not a character flaw.

It's what happens when a group commits to one story early. Nobody ever makes it defend that story against a competitor. The story wins by default, not by merit.

The 5 Whys doesn't fail because people ask bad questions. It fails because it never asks them to consider a second answer.

The Alternative: Tracking Multiple Working Hypotheses

Here's the fix: stop looking for the one root cause. Track several at once, side by side. Let the evidence knock most of them out.

Nothing new here, either. People have been doing a version of this for over a century. Not in factories. In dirt.

Paper trail, if you want it: a geologist named Chamberlin wrote it up in 1890.

It rarely survives contact with a real shop floor, though. Holding four or five competing explanations in your head is a lot to ask of anyone. Especially while you're also running a floor.

Chamberlin's method has stuck around this long for one reason. The problem never went away.

People fall for the first explanation that sounds right. Then they stop looking. It happens in a conference room as easily as it happened in a field survey a century ago.

This is the overhead AI is built to carry:

●       Holds the structure so you don't have to keep it all in your head.

●       Pulls evidence from your data faster than you could by hand.

●       Keeps every hypothesis honest instead of just the one you already like.

What it should never do is pick a winner for you.

Building an AI-Powered Root Cause Hypothesis Matrix

Instead of a single chain of whys, build a matrix. Six pieces of information, for every hypothesis:

●       The hypothesis itself. Stated as a guess, not a conclusion.

●       Evidence supporting it.

●       Evidence contradicting it.

●       Evidence you don't have yet.

●       A specific test that could kill it, or move it forward.

●       What you'd expect to see, if it's true.

Here's what that looks like, applied to the Line 4 lead. Failures clustered around low staffing, product family B, and the first ninety minutes after a changeover.

Don't accept that combination as the answer. Treat it as one hypothesis among several. Test them side by side:

Hypothesis 1: New relief operator is undertrained for solo coverage during breaks

●       Supporting evidence: Failures cluster during known relief windows on two of the five weeks checked

●       Contradicting evidence: Failures also occur when the regular operator runs the line solo, at a similar rate

●       Missing evidence: The relief operator's training completion date and prior line assignments

●       Test: Cross-reference every flagged failure against the shift roster for that exact hour

●       Expected result if true: The failure rate would drop sharply once relief-operator shifts are excluded

Hypothesis 2: Product family B runs closer to tolerance, so normal variation trips more faults

●       Supporting evidence: Family B shows a higher fault rate across all shifts, not just the flagged windows

●       Contradicting evidence: Line 2 runs the same product family without an elevated fault rate

●       Missing evidence: Machine-level calibration logs for both lines on family B runs

●       Test: Compare family B fault rates across every line that runs it, not just Line 4

●       Expected result if true: Other lines running family B would show a similar, if smaller, elevated rate

Hypothesis 3: The changeover procedure leaves a setting unconfirmed, and it drifts over the next ninety minutes

●       Supporting evidence: Every flagged failure falls within ninety minutes of a changeover

●       Contradicting evidence: Not every changeover is followed by a failure, only those paired with low staffing

●       Missing evidence: Setup signoff logs for flagged versus unflagged changeovers

●       Test: Pull signoff records for ten flagged and ten unflagged changeovers and compare

●       Expected result if true: The signoff step would be skipped or rushed disproportionately in the flagged group

Notice what this does that a 5 Whys chain can't. None of these three hypotheses has been declared the winner. Each one has a specific, falsifiable test attached to it.

The contradicting evidence is doing real work in every one of them. It keeps everyone honest. It forces you to write down what doesn't fit, before you fall for the story that does.

The One Rule That Keeps This Honest: Falsifiability

Every hypothesis in that matrix earns its spot because it can be proven wrong. That's the whole mechanism. Not a minor detail.

If you can't picture the test that would kill a hypothesis, it isn't a lead. It's a guess wearing a lab coat. There's a technical term for a claim you can disprove: falsifiability.

Use that as your filter, every time AI hands you a list of explanations. Falsifiability is the difference between a real investigative lead and a story that sounds good.

If a hypothesis doesn't come with a test that could kill it, send it back. Ask what evidence would prove it wrong. Not just what evidence supports it.

That one question protects you from a bad root cause better than any amount of extra data ever will.

The Prompt: Build Your Own Hypothesis Matrix

This works best with a tool that can read files and hold a structured conversation, not just answer a single question: Claude, ChatGPT with data analysis enabled, or similar. Take the strongest lead from your last investigation, along with whatever supporting data you have on it. Then adapt this prompt to your situation.

1. State the Lead, Not a Conclusion

●       Describe the pattern you found as a hypothesis to investigate, not a confirmed cause.

●       Include the specific conditions involved and the data window it came from.

2. Ask for Competing Explanations, Not Confirmation

●       Instruct it to generate at least two or three alternative hypotheses that could also explain the same pattern.

●       Include at least one that has nothing to do with your leading theory.

●       Example instruction: “Don't just confirm my hypothesis. Give me at least three plausible alternative explanations for this same failure pattern, including at least one I haven't considered.”

3. Build the Six-Column Matrix

●       Ask it to lay out each hypothesis with supporting evidence, contradicting evidence, and missing evidence.

●       Add a specific test for each one, and the expected result if that hypothesis is true.

4. Force a Falsifiability Check on Every Row

●       Ask it to flag any hypothesis that doesn't have a test capable of disproving it.

●       Either sharpen that test, or drop the hypothesis from the matrix.

5. Rank the Hypotheses by Testability, Not Plausibility

●       Test the cheapest hypothesis first, even if a different one feels more likely.

●       Cheap tests clear the field fast.

6. Keep the Matrix Alive as Evidence Comes In

●       After you run a test, feed the result back in and ask it to update every row, not just the one you tested.

●       Evidence against one hypothesis is often evidence for or against the others too.

Run that loop and the matrix does the thinking for you. It still can't do the part that only happens standing on the floor.

What AI Can't Decide for You

AI can build the matrix. It can populate the evidence columns from your data. It can flag a weak test before you waste a week on it.

It cannot walk out to Line 4 and watch the changeover happen. It cannot do any of the following:

●       Pull a setup sheet out of a drawer.

●       Interview the relief operator.

●       Notice that the fault light flickers a half-second before the line stops.

It also can't smell burnt gearbox oil. That one still matters.

Every test in that matrix must still be run by a person who knows the floor.

What changes is what you're testing. You're not chasing the first explanation that felt right in a meeting anymore. You're running a cheap test against a hypothesis that already survived an attempt to kill it.

That's a different use of your time on the floor. The walking and watching part is still entirely yours.

Next Up in the Series

Say your test confirms the changeover hypothesis. You're ready to fix the signoff step in the setup procedure.

Before you touch it, you need to know what that fix does downstream.

Next up, we use AI to model constraint migration. Here's what happens to WIP, labor, and dock capacity when your fix at Line 4 shifts the bottleneck somewhere you weren't watching. Link: You Didn't Fix the Bottleneck: How AI Predicts Where it Moves Next.

Subscribe to The Daily Constraint. The next one lands in your inbox the day it goes live.

Build a matrix this week? Comment below. Tell me which hypothesis survived. Not which one you expected to win.