CASE FILE

How an Error Happens That No One Made

Analyzing System Failure

Jak vzniká chyba, kterou nikdo neudělal
Jiný Kontext editorial illustrationAnalýza systémového selhání
Listen
00:00/00:00
1.00 ×
Ready

At the eighty-second second of the flight, a piece of insulating foam broke away from the Space Shuttle Columbia’s external tank. On the recording, it was not an explosion, a fire, or a dramatic warning. Just a pale fragment, a brief movement across the frame, and an impact with the left wing. The mission continued. The crew worked in orbit. On Earth, meanwhile, a process began in which a technical signal gradually became a question, the question an estimate, and the estimate a reason to do nothing further.

Chapter 01

The Foam That Did Not Look Like a Verdict

During the mission, NASA engineers made three requests for high-resolution images of the shuttle. The requests did not pass through the proper channels. One reached the relevant people at the Department of Defense, but NASA canceled it approximately ninety minutes later. The investigative board later described a flawed analysis of the possible damage, low concern among senior management, unclear communication, and a weak role for safety structures. It also noted an important uncertainty: even in hindsight, it was not certain that the requested images would actually have shown the damage.[1]

On February 1, 2003, Columbia broke apart during reentry. Hot gases entered through a breach in the leading edge of the left wing, compromised its structure, and aerodynamic forces completed the destruction. All seven crew members died. The physical mechanism was specific: approximately 0.77 kilograms of foam struck the wing 81.9 seconds after launch. The Columbia Accident Investigation Board report, however, gave equal weight to organizational causes—schedule pressure, budget constraints, communication barriers, informal decision-making channels, and reliance on past success instead of sufficiently grounded engineering judgment.[1]

From today’s perspective, the sequence of events looks almost unbearably straightforward: something hit the wing, engineers wanted images, the images were not taken, and the shuttle broke apart. Once we know the ending, every earlier moment appears to point toward it.

But the people making decisions at the time did not know the ending. They saw a fragment, not a breach. They had models, not certainty. They knew foam had come off before and missions had returned safely. They worked in an organization that had to protect the crew, keep the program running, distinguish serious signals from hundreds of less significant anomalies, and make decisions with limited data—all at once.

And this is where the whole problem changes. A catastrophe need not be the product of one insane decision. It can emerge as the sum of steps, each defensible within its own small slice of reality.
A Linear Story and a Systemic Picture of an Accident The left side shows a simple chain with one culprit; the right, a network of conditions and layers of defense. LINEAR STORY SYSTEMIC PICTURE Looks for one wrong step and one author. Traces how conditions and defenses combined. Someone made a mistake “bad decision” An accident occurred one cause → one consequence We remove the culprit and assume the problem is fixed Understandable. Reassuring. Often incomplete. Incomplete datauncertain signal Past successfalse reassurance Performance pressuretime, cost, operations Distributed authoritygaps between departments Weak feedbacknear miss without learning Degraded defensescontrols only on paper FAILURE emerges through combination
Graphic 01 One Arrow Against a Network of ConditionsA linear story works for simple mechanisms. In a complex institution, however, it often mistakes the last visible step for a complete explanation. Editorial illustration—a simplified model.
Chapter 02

The Temptation to Find a Single Culprit

After an accident, a wrongful verdict, or a data breach, the same question almost always appears: Who messed up? It is practical because an institution needs to act. It matters legally because responsibility cannot be dissolved into an abstract “system.” And it is psychologically appealing because it turns chaos into a story with a character, a decision, and a consequence.

The problem begins when we mistake the question of responsibility for the question of causation. British psychologist James Reason distinguished a “personal” view, focused on an individual’s mistake, from a “systemic” view, which examines working conditions and layers of defense. In his famous model, safety barriers resemble slices of cheese with holes: no defense is perfect, but their weak spots usually do not line up. An accident occurs when they briefly do.[2]

Richard Cook went further. In complex systems, he argued, a catastrophe is often the result of multiple failures, each necessary but none sufficient on its own. Systems are full of small defects, ambiguities, improvisations, and temporary fixes; they function nevertheless because they have reserves and because people continually compensate for their imperfections. Searching for a single “root cause” after an accident can therefore be misleading. We often find only the last link, the one close enough to the consequence to be easy to name.[3]

The immediate cause explains what happened at the end. Systemic analysis asks why a path to that ending was able to open at all.
Distinguishing Mechanism from Conditions

This does not mean that nothing is anyone’s fault. It means only that firing an operator, punishing a technician, or replacing a manager may not remove the mechanism that shaped their decisions. A new person can enter the same shift, the same software, the same priority system, and the same fog of information—and eventually do the same thing.

The individual is visible. Interfaces between departments, escalation rules, ways of reporting uncertainty, and long-tolerated deviations are not. That is why, after an accident, we so readily punish the former and find it so difficult to repair the latter.

Chapter 03

Reasonable Steps, an Unreasonable Whole

A decision that looks incomprehensible after a tragedy may have been locally rational beforehand. This concept, associated with the work of David Woods and Richard Cook, does not say that people made the right decisions. It says their actions must be reconstructed from the information, goals, constraints, and pressures they had at that moment—not from the outcome that only we know.[4]

Imagine an ordinary chain. A technician records a deviation that has appeared several times without consequences. An engineer receives incomplete data and a model that does not anticipate serious damage. A manager sees a technical summary, not the original record, and must decide whether to activate an expensive emergency procedure. Senior management hears that there is no formal “request,” only an unverified concern. No one has to lie. No one has to be indifferent. Each person is simply working with a different version of the problem.

Limited informationno one sees the whole system
+
Conflicting goalssafety, time, capacity, cost
+
Weak couplingthe signal loses meaning between layers

Once we know the outcome, hindsight bias enters the picture. Events that had been competing with dozens of other possibilities are retrospectively arranged into a single inevitable trajectory. The warning seems clearer, the alternative more available, and the wrong choice more absurd than it was in real time. Hindsight bias is not an excuse; it is an obstacle to investigation. If we do not filter it out, instead of analyzing the conditions of decision-making we will analyze a caricature of a person who supposedly had to know what we know only after the accident.

Jens Rasmussen described organizations as systems pushed in several directions at once: toward greater efficiency, lower costs, and a lighter workload, but also toward the boundary of acceptable risk. Operations do not have to approach that boundary in one conscious leap. They move toward it millimeter by millimeter, through small adaptations that look reasonable individually.[5]

An important distinction

Local rationality is not a moral acquittal. It is an analytical tool. It lets us ask a more precise question: not “How could anyone do something so stupid?” but “What information and incentives made this step seem acceptable at the time?”

Chapter 04

When a Warning Learns to Be Normal

An organization’s most dangerous experience may not be failure. Sometimes it is repeated success despite breaching a safety margin.

Sociologist Diane Vaughan, studying the decisions that preceded the Challenger disaster, described the “normalization of deviance.” A repeatedly observed anomaly did not become a reason to stop; because it had not led to catastrophe in the past, it was gradually incorporated into the picture of normal operations. In her account, this was neither a simple conspiracy nor a one-time disregard for risk. It was a gradual descent into a judgment that no longer felt deviant within that culture.[6]

The mechanism is deceptive. The first deviation causes concern. The second resembles the first, but because the first ended well, it is slightly less alarming. By the tenth, the organization no longer reads the repetition as ten warnings, but as ten pieces of evidence that the situation is manageable. The absence of catastrophe turns into confirmation of safety.

Yet a safe return can mean two entirely different things: either the margin was genuinely sufficient, or the system survived this time because of circumstances it did not control. Without measurement, the two possibilities cannot be reliably distinguished.

The Normalization-of-Deviance Loop Five steps show how a repeated deviation without immediate harm weakens perceived risk and expands the accepted norm. THE BOUNDARY SHIFTS WITHOUT LOOKING LIKE IT 1 — DEVIATION A rule or margin is exceeded the situation feels unusual 2 — NO HARM Nothing happens this time luck looks like evidence 3 — NEW INTERPRETATION The risk is judged to be smaller past success changes expectations 4 — NEW NORM The deviation is repeated it no longer attracts the same attention 5 — GROWING EXPOSURE the next cycle begins from a shifted boundary
Graphic 02 How an Exception Becomes the StandardSuccess without harm does not, by itself, reveal whether the deviation was safe or merely failed to produce a consequence this time. Editorial illustration based on Diane Vaughan’s concept of normalization of deviance.

Another paradox enters this process. Front-line workers often keep a poorly designed system running precisely by improvising. They work around a broken form, supply missing information by phone, catch a colleague’s error, and repair inconsistent data. Their adaptations create real-time safety. At the same time, they conceal from leadership how fragile operations actually are. The organization sees the result—a completed task—not the hundreds of small interventions without which it would fall apart.[3]

Then one person is absent, two safeguards fail at the same time, or an unusual combination of circumstances appears. What looks like a sudden collapse is actually the moment when people can no longer compensate for defects that have been there for a long time.

Chapter 05

How Information Gets Lost Without Anyone Concealing It

In organizations, information is almost never passed on in its original form. It must be sorted, translated, shortened, and placed into a format the next level can understand. A technician describes an anomaly. An engineer turns it into an estimate. A manager converts the estimate into a risk category. Management compares that category with the schedule, budget, and formal rules.

Every translation is necessary. A director cannot monitor every sensor, and a judge cannot conduct the entire investigation again. At the same time, each translation loses part of the context—especially uncertainty, contradictions, and weak signals that are difficult to fit into a single box.

This is how organizational ignorance can arise without a single lie. It is not that the information was never sent “upward.” It arrives in a form that no longer raises the same questions. In the Columbia case, a technical concern became, as it passed through the structure, the absence of a formal request for imaging. The difference was not created by new facts. It was created by the way the institution decided what counted as sufficient reason to act.[1]

How a Warning Signal Loses Strength as It Passes Through an Organization Five stages show the transformation of a raw signal into a management decision, while detail and uncertainty gradually disappear. A SIGNAL PASSES THROUGH FIVE TRANSLATIONS 1 RAW SIGNAL record, photograph, deviation, doubt maximum detail 2 TECHNICAL INTERPRETATION model, probability, comparison with the past uncertainty is named 3 MANAGEMENT SUMMARY risk category, recommendation, urgency some contradictions disappear 4 FORMAL STATUS request / input / no action context becomes a checkbox 5 DECISION act, defer, continue minimum of the original detail INFORMATION COMPRESSION Translation is necessary. The danger arises when the process removes, along with the noise, the uncertainty that should have changed the decision.
Graphic 03 A warning does not have to be silenced. It only has to be translated into a weaker language.The diagram does not describe a specific institution; it shows the general problem of information compression between expert, management, and decision-making levels. Editorial illustration—a simplified model.

The same problem arises in the opposite direction. Leadership issues a general rule that encounters physical reality, staff shortages, or an exception the design did not anticipate on the front line. Employees create a practical shortcut. The shortcut is not recorded because formally it should not exist. Leadership therefore continues to believe that the process works as documented.

The organization then lives in two versions of itself: the official one, where steps are clear and responsibilities assigned, and the operational one, where work gets completed through improvisation. A catastrophe often exposes the difference only when it can no longer be bridged.

Chapter 06

A Miscarriage of Justice Without a Single Moment of Error

In a miscarriage of justice, we tend to look for a single turning point: a false confession, a mistaken identification, a manipulated expert opinion, or the deliberate concealment of evidence. Such turning points exist. But many cases become persuasive only when several weak elements begin to confirm one another.

A witness sees the perpetrator only briefly, under stress and in poor light. Later, the witness chooses from photographs. An investigator who knows the suspect may unconsciously react to hesitation. After making the choice, the witness remembers the face with more certainty than before. The file contains the identification, not the entire process by which that certainty arose. Another investigator then reads it as an established fact. New evidence is assessed in its light.

The U.S. National Research Council warned that perception, memory, and subjective certainty are malleable. Memories are reconstructed and updated during encoding and retrieval, and can be distorted without the witness realizing it. Accuracy is influenced by the length of observation, distance, lighting, stress, and the identification procedure itself. The report therefore recommended, among other measures, double-blind lineups, standardized instructions, verbatim recording of the initial level of confidence, and video recording of the procedure.[7]

None of these steps creates a verdict on its own. A verdict emerges from their connections: witness memory, the wording of a question, the selection of a suspect, the interpretation of a forensic finding, prosecutorial decisions, the capacity of the defense, and the rules by which a court admits or weighs evidence. A weak piece of evidence can reinforce another weak piece because both arise from the same original assumption. A circle forms that looks from the outside like several independent confirmations.

Factors in U.S. Exonerations in 2024 The horizontal bar chart shows the percentages of contributing factors among the 147 exonerations recorded by the National Registry of Exonerations in 2024. CONTRIBUTING FACTORS IN 147 U.S. EXONERATIONS Cases recorded as exonerations that occurred in the United States in 2024 0 %20 %40 %60 %80 % False testimony / false accusation Official misconduct / improper conduct by officials Inadequate legal defense False / misleading forensic evidence Mistaken eyewitness identification False confession 72 % · 106 71 % · 104 33 % · 48 29 % · 42 26 % · 38 15 % · 22 CATEGORIES OVERLAP. A single case may contain several factors; the total is therefore not 100%. These are known U.S. exonerations, not an estimate of the frequency of all miscarriages of justice or all convictions.
Graphic 04 Wrongful Convictions Usually Have More Than One IngredientThe National Registry of Exonerations recorded 147 exonerations that occurred in the United States in 2024. The listed factors overlap, and the database captures only cases that were uncovered. It cannot be used to infer the share of wrongful convictions among all convictions.[8]
Data source: The National Registry of Exonerations, 2024 Annual Report, April 2, 2025.

The numbers show something else that complicates this article’s title. In 71 percent of these exonerations, the registry recorded official misconduct or improper conduct; in many cases, several forms of conduct were present at once. This is not a picture of a world in which everyone is merely an innocent cog. It is a picture of a world where systemic weaknesses, human errors, professional failures, and sometimes deliberate abuse of power overlap.[8]

That is why precision matters: “an error that no one made” does not mean that no one did anything. It describes a situation whose outcome cannot honestly be explained by a single act or a single person. Sometimes we really do find a lie, manipulation, or gross negligence in the chain. Even then, it may not explain the entire chain.

Chapter 07

The System Is Not an Alibi

Systems thinking has its own danger. It can turn into a fog in which individual responsibility disappears. If everything was caused by “culture,” “process,” or “complexity,” it may eventually seem that no one is responsible for anything. That would be just as inaccurate as searching for a single scapegoat.

The approach known as just culture seeks to distinguish different types of conduct: an ordinary human error, a risky shortcut that has become the norm, and a knowingly reckless violation of rules. The response should not be based only on the magnitude of the outcome. The same dangerous step may end harmlessly once and tragically the next time; the moral quality of the conduct has not changed. Conversely, an unintentional error in a poorly designed interface does not become blameworthy merely because it had a major impact this time.[9]

Mechanism

What physically or procedurally caused the outcome? Missing bolts, an incorrect data point, a medication administered incorrectly, a mistaken identification.

Conditions

What environment made the error more likely? An unclear procedure, workload, a missing check, conflicting goals, a poor interface.

Conduct

Was it an understandable error, a normalized shortcut, a deliberate violation, or deception? This is where personal responsibility is determined.

System governance

Who knew about the recurring risk, who had the authority to remove it, and how did the organization respond to earlier warnings?

The final U.S. NTSB report on Alaska Airlines Flight 1282 illustrates this well. On January 5, 2024, an emergency-exit door plug separated from a Boeing 737-9 during climb; four retaining bolts had not been installed. The report issued in July 2025 did not identify only the immediate technical mechanism. It identified Boeing’s failure to provide adequate training, instructions, and oversight of the removal and reinstallation process as the probable cause. It identified ineffective FAA oversight and audit planning, which failed to catch recurring systemic discrepancies, as a contributing factor.[10]

The missing bolts were real. They remained out of the structure during reassembly. Yet an investigation would end too soon if it merely found the last worker at the aircraft. The question was not only who performed the task, but why the process did not ensure that the opening was documented, why it did not trigger an independent inspection, and why oversight did not turn recurring discrepancies into effective corrective action.

What Systemic Analysis Must Not Do

It must not confuse understanding with excuse. Explaining why conduct arose in a given environment does not mean denying a deliberate lie, gross negligence, or abuse of authority. It means refusing the comforting idea that punishing a person automatically repairs the conditions that created the problem.

Chapter 08

Institutions That Know How to Doubt

A reliable organization is not one in which people never make mistakes. No such organization exists. A more reliable one is built to expect incomplete information, conflicting goals, and changing operations—and to create more opportunities to catch, challenge, or correct an error before it combines with others.

Research on high-reliability organizations often summarizes their approach in five principles: a constant preoccupation with failure, reluctance to simplify interpretations, sensitivity to what is actually happening in operations, deference to expertise regardless of hierarchy, and the ability to restore a safe state when something goes wrong. The evidence on implementing these principles is limited, and no set of slogans creates safety by itself. They matter only when they become everyday decision-making habits.[11]

01 Preserve the weak signal

Record uncertainty, disagreement, and the original data alongside the conclusion. The summary must not be the only surviving version of the problem.

02 Separate safety from performance pressure

The person responsible for risk needs the authority to escalate a problem, even when doing so disrupts a deadline, budget, or reputation.

03 Record events without harm

An event without harm is not a reason to forget. It is an inexpensive view of a trajectory that may not be stopped next time.

04 Give expertise a voice

In technical uncertainty, the decisive voice should not automatically belong to the highest-ranking person, but to the person with the most relevant expertise.

05 Design a route back

Safety is not only prevention. It also includes the ability to stop a process, verify the state, reverse a change, correct an error, or save lives.

06 Investigate without a preselected culprit

First reconstruct the work as it actually happened. Only then distinguish error, risky adaptation, and blameworthy conduct.

Such institutions may appear less confident from the outside. They produce more reports, more dissenting views, and sometimes more recorded errors. That does not necessarily mean they are less safe. It may mean they see what another organization is still suppressing. A low number of reported problems is good news only if we know the system is genuinely looking for problems and people can report them without fear of automatic punishment.

After the Columbia accident, the CAIB recommended more than technical changes. It called for an independent technical authority, stronger safety structures, better imaging, a realistic schedule in relation to resources, and the ability to inspect and, if necessary, repair the shuttle in orbit. In other words, building a stronger wing was not enough. The way uncertainty reached decisions had to change, as did the options that remained when prevention failed.[1]

The same logic applies in justice. A double-blind lineup does not rely on an investigator’s avoiding influence on the witness; it removes the possibility of unconscious influence from the procedure’s design. Video recording does not guarantee a correct identification, but it preserves a process that would otherwise be reduced to its result. A verbatim record of the initial level of confidence prevents later confidence from rewriting the original hesitation.[7]

The best safeguard is therefore often not a better person. It is an arrangement that does not require a person to be infallible.

Conclusion

The Empty Space Between Decisions

Let us return to the image from the eighty-second second. A pale fragment breaks away from the tank, disappears near the left wing, and the shuttle continues upward. From our perspective, the image is already filled with the future. We see reentry, the vehicle breaking apart, and seven lives that will end sixteen days later.

The people at that moment saw something else: a familiar type of anomaly, an incomplete record, and a problem that had to be placed among many others. Some wanted more data. Others trusted the model. Still others heard that there was no formal request. The catastrophe was not created by an absence of intelligence. It was created by the organization of intelligence—the way information, authority, doubt, and opportunities to intervene were distributed.

The question “Who made the mistake?” therefore remains important, but it is not enough. It must be joined by others: Who saw which part? What was treated as evidence? Which past successes reduced sensitivity to risk? Where was uncertainty lost? Who had responsibility without authority, and who had authority without an adequate picture? Which safeguard existed only in the documentation, and which was actually replaced by human improvisation?

Only these questions can distinguish accidental error, systemic blindness, a normalized shortcut, and deliberate failure. And only this distinction makes it possible both to assign responsibility and to reduce the likelihood of recurrence.

The greatest institutional error is not always one bad decision. Sometimes it is the empty space between several decisions, each of which looked reasonable on its own.

Sources and literature

Sources and further reading

  1. Columbia Accident Investigation Board / Congressional Research Service — Columbia Accident Investigation Board Report, Volume I; souhrn CRS (2003)

    Supports the technical mechanism of the accident, the imaging requests, organizational causes, and post-accident recommendations.

    https://ntrs.nasa.gov/citations/20030093634
    NASA / CRS: Synopsis of the CAIB Report (PDF)
  2. James Reason — Human error: models and management (BMJ, 2000)

    Source for the distinction between personal and systemic approaches and the model of multiple imperfect layers of defense.

    https://doi.org/10.1136/bmj.320.7237.768
  3. Richard I. Cook — How Complex Systems Fail (2000)

    Supports claims about multiple simultaneously necessary failures, degraded-mode operation, the limits of a single “root cause,” and the adaptive role of people.

    https://www.adaptivecapacitylabs.com/HowComplexSystemsFail.pdf
  4. David D. Woods & Richard I. Cook — Perspectives on Human Error: Hindsight Biases and Local Rationality (1999)

    Expert framework for local rationality, hindsight bias, conflicting goals, and reconstructing decisions under their original conditions.

    https://how.complexsystems.fail/citations/Perspectives_on_Human_Error.pdf
  5. Jens Rasmussen — Risk management in a dynamic society: a modelling problem (Safety Science, 1997)

    Supports the model of operations gradually moving toward the boundary of acceptable performance and risk under simultaneous pressures.

    https://orbit.dtu.dk/files/158016663/SAFESCI.PDF
  6. Diane Vaughan — The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA (1996; expanded edition 2016)

    Foundational sociological work on the normalization of deviance and the culturally conditioned acceptance of repeated anomalies.

    https://press.uchicago.edu/ucp/books/book/chicago/C/bo22781921.html
  7. National Research Council — Identifying the Culprit: Assessing Eyewitness Identification (National Academies Press, 2014)

    Supports the account of the limits of perception and memory and recommendations for blinded procedures, standardized instructions, confidence recording, and video-recorded lineups.

    https://doi.org/10.17226/18891
  8. The National Registry of Exonerations — 2024 Annual Report (April 2, 2025)

    Source for the number of 147 U.S. exonerations in 2024 and the shares of overlapping contributing factors used in the chart.

    https://exonerationregistry.org/sites/exonerationregistry.org/files/documents/2024_Annual_Report.pdf
  9. Agency for Healthcare Research and Quality, PSNet — Culture of Safety (continuously updated expert overview)

    Supports the “just culture” framework and the distinction between human error, risky conduct, and reckless violation.

    https://psnet.ahrq.gov/primer/culture-safety
  10. National Transportation Safety Board — DCA24MA063, Alaska Airlines Flight 1282, Final Report (July 11, 2025)

    Supports the facts about missing retaining bolts, inadequate Boeing training, instructions, and oversight, as well as the contributing role of FAA oversight.

    https://data.ntsb.gov/carol-repgen/api/Aviation/ReportMain/GenerateNewestReport/193617/pdf
  11. Agency for Healthcare Research and Quality, PSNet — High Reliability Organization Principles and Patient Safety (2025)

    Expert discussion of the five principles of high reliability, their practical implementation, and the limits of the available evidence.

    https://psnet.ahrq.gov/perspective/high-reliability-organization-hro-principles-and-patient-safety
Discussion

Comments

Have a view or an additional source? Add a comment.

Discussion is not active yet.