The outage postmortem has a familiar shape by the third time you’ve run one for the same incident. Someone explains what happened. The explanation is specific, technical, and reasonable, a bad deploy, a missed alert, a tired engineer at 2 a.m. The room nods. Someone writes an action item. Six weeks later, a slightly different version of the same failure takes down the same system, and the postmortem gets scheduled again.
Nothing in that cycle was dishonest. The explanations were all true. They just weren’t the cause, they were the last visible link in a chain nobody bothered to trace further back. Fixing the last link doesn’t break the chain. It just moves the failure a few weeks down the calendar.
A Question Toyota Took More Seriously Than Most
The discipline for tracing that chain properly has a name simple enough to undersell it: the 5 Whys. Ask why the problem happened. Ask why about that answer. Keep going. Sakichi Toyoda built the habit into what would become the Toyota Production System in the 1930s, and Taiichi Ohno, the engineer most responsible for turning Toyota’s shop floor into a discipline other companies would spend decades trying to copy, called it plainly “the basis of Toyota’s scientific approach”: by repeating why five times, he wrote, “the nature of the problem as well as its solution becomes clear.”1
The number five was never the point. What Ohno was actually describing was a stopping rule, and it’s a better one than “we found an explanation.” The version used in federal quality-improvement guidance states it cleanly: before you commit to a fix, ask whether the problem would still happen again if this particular answer got corrected.2 If the honest answer is yes, the chain isn’t finished. You’re still standing on a symptom, dressed up as a cause because it was the first one specific enough to write down.
Business coach David Jenyns, who has run this exercise with small operating teams for years, puts a number on the discipline: don’t let a group decide on a fix until it has asked “why” at least four times.3 The first two answers almost always land on a person, a rushed engineer, a distracted rep, a vendor who missed a deadline. Push past them, and the chain usually resolves into one of three things instead: a process nobody documented, a template or safeguard that was quietly missing, or a step in the workflow with no one clearly responsible for catching a failure before it reached a customer.3
a Team Is Allowed to Pick a Fix
State the Problem Before You Chase the Cause
The chain only works if it starts somewhere real. “This keeps taking too long” is a complaint, not a problem statement, because there’s nothing in it to ask “why” about yet. A problem is a specific, checkable gap between what’s happening and what should be happening, stated as a fact the room can verify, not a feeling about how the team performs. “Deploys are averaging three rollbacks a month, up from near zero a year ago” gives the first why something to answer. “Nobody takes releases seriously anymore” gives it nothing but a person to blame, which is usually where these sessions quietly go wrong before the first real question ever gets asked.
Emotion does specific, avoidable damage to a session like this. A statement like “this team keeps dropping the ball” points the first why at a person before anyone has asked a single question, and it swaps a fact the room could trace for a feeling the room ends up arguing about instead. The fix isn’t stoicism, it’s precision: if a word in the problem statement describes how you feel about the failure rather than something you could point to on a dashboard, cut it before the session starts.
Where the Discipline Breaks
The 5 Whys has a documented weakness, and it’s worth knowing before you lean on the technique as if it were beyond question. Teruyuki Minoura, a former managing director at Toyota, identified several ways the method fails in practice: investigators tend to stop at symptoms rather than pushing to a real lower-level cause, the questioning has no structured way to surface answers outside what the investigator already knows, and, most consequentially, the technique tends to isolate a single root cause when a real failure often has two or three converging on it at once. Different people running the identical session, he noted, frequently arrive at different “root causes” from the same starting problem.1
Patient-safety researcher Alan J. Card pushed the critique further in a 2017 paper in BMJ Quality & Safety, arguing that the fixed five-level depth is an arbitrary number that rarely lines up with where an actual root cause sits, and noting that the technique was originally built for Toyota’s product-development troubleshooting rather than formal root cause analysis, a distinction lost somewhere in its decades of borrowed use across industries that never questioned the fit. His recommended complement is a branching method, a fishbone diagram, that can hold more than one contributing cause at a time instead of collapsing everything onto a single chain.4
None of that makes the discipline useless, it makes it incomplete on its own. The practical response for a team that relies on it is to treat any single 5 Whys chain as a hypothesis, not a verdict. If the group can’t agree on the answers at each step, or if the fix that came out of the session doesn’t actually stop the failure from recurring, that’s not evidence the method failed, it’s evidence there was a second contributing chain the first session never touched. Run it again with someone who saw a different part of the system, and pair whatever root cause the team does settle on with a written change, a checklist, an updated runbook, an actual owner, not just a verbal agreement that dissolves the moment the meeting ends.
The Fix Is Usually Smaller Than It Feels
What makes this discipline hard to sustain isn’t the logic, it’s the patience. Four or five questions deep, the answer on the table is rarely dramatic. It’s a missing checklist, an alert that goes to the wrong channel, a handoff nobody formally owns. That can feel like an anticlimax after a failure serious enough to justify the meeting, which is exactly why so many rooms stop early and reach for something that feels proportionate instead, a reorg, a new tool, a stern conversation. The failure was never proportionate to begin with. It was one small, specific gap, repeating quietly for months because nobody had asked about it four times in a row.
Key Takeaways
Key Frameworks:
- 5 Whys: repeatedly asking “why” about the previous answer until the chain reaches a process, template, or ownership gap instead of a person or a one-off event.
- The recurrence test: before committing to a fix, asking whether the problem would still happen again if this particular answer got corrected.
Try It: Pick one problem that has happened more than once on your team in the last month. Write it as a single checkable sentence, with no names and no adjectives about effort or attitude. Ask “why” four times with whoever is closest to the work, and write down one fix, one owner, and the date you’ll check whether it actually stopped.
-
“Five whys,” Wikipedia, https://en.wikipedia.org/wiki/Five_whys ↩↩
-
CMS, “Five Whys Tool for Root Cause Analysis,” https://www.cms.gov/medicare/provider-enrollment-and-certification/qapi/downloads/fivewhys.pdf ↩
-
David Jenyns, “5-Whys Problem Solving for Small Business,” https://www.davidjenyns.com/5-whys-problem-solving-small-business/ ↩↩
-
Alan J. Card, “The problem with ‘5 whys’,” BMJ Quality & Safety 26, no. 8 (2017): 671-677, https://doi.org/10.1136/bmjqs-2016-005849 ↩