An employee selects the wrong component. A technician records the wrong value. An operator skips a step. A reviewer misses an error. A required check is not completed.
The investigation opens. A few days later, the conclusion reads:
Root Cause: Human Error. Corrective Action: Retrain the employee. CAPA closed.
And six months later, something very similar happens again.
That pattern is one of the clearest signs that an investigation stopped at the event instead of finding the conditions that allowed the event to happen.
“Human error” can accurately describe what an individual did. But it often does not explain why the quality system allowed one ordinary mistake to become a product, process, data, or compliance failure.
A stronger investigation asks: Why was this mistake possible? Why wasn't it prevented? Why wasn't it detected earlier? Could the same conditions affect someone else tomorrow?
Those questions are where root cause analysis begins.
The Person Closest to the Failure Is Not Automatically the Root Cause
Start here
The person who made the final mistake may be the last link in the causal chain, not the first.
Consider a wrong-component event. The employee physically selected the wrong item. That is factual. But the investigation may also find that:
- Two components had nearly identical packaging.
- Part numbers differed by only two digits.
- The materials were stored next to each other.
- The procedure relied entirely on visual verification.
- There was no barcode or system verification.
- A previous near-miss involving the same components had already occurred.
- Production schedules encouraged rapid changeover.
- No independent check existed before use.
Saying “the operator selected the wrong part” identifies what happened. It does not explain why the process was designed so that one selection mistake could pass through every available control.
That distinction is the heart of effective root cause analysis.
FDA Has Made This Point Explicitly
A 2026 FDA Warning Letter Provides a Clear Example
In a May 27, 2026 warning letter to pharmaceutical manufacturer Pharmathen International S.A., FDA addressed recurring sterility and media-fill failures in which operator error had contributed significantly.[1]
FDA did not simply accept “operator error” as the end of the investigation. The agency stated that the firm's investigation:
FDA / Regulatory Example
“did not include in-depth analysis of the conditions that create the circumstances for human error.”
— U.S. Food and Drug Administration, Warning Letter to Pharmathen International S.A., May 27, 2026
FDA went on to state that an in-depth root-cause analysis was needed to identify effective CAPA, including consideration of design remediation. The agency also criticized the firm's proposed CAPAs because they lacked sufficient detail about how effectiveness would be evaluated.
This is an important distinction. The regulatory concern was not merely that employees had made errors. The concern was that the investigation had not adequately examined why those errors were recurring or what conditions within the operation made them possible.
Source: U.S. Food and Drug Administration, Pharmathen International S.A. — Warning Letter 320-26-80, May 27, 2026.
Scope note
This warning letter concerns pharmaceutical CGMP requirements under 21 CFR Parts 210 and 211. It should not be read to mean that every statement in the letter establishes a universal requirement for ISO 9001, ISO/IEC 17025, medical-device, laboratory, or other QMS environments. It is cited here as a current FDA enforcement example demonstrating why investigations may need to look beyond the immediate human action.
“Human Error” Is Often a Stopping Point Disguised as a Cause
When an investigation concludes “human error,” ask what would happen if the employee were replaced tomorrow. Would the next appropriately trained employee face:
- The same confusing screen?
- The same similar labels?
- The same unclear instruction?
- The same manual transcription?
- The same workload?
- The same missing verification?
- The same equipment limitation?
- The same opportunity to make the same mistake?
If the answer is yes, then replacing or retraining the employee has not removed the condition that produced the failure.
That does not mean the employee bears no responsibility. It means accountability and causality are different questions. An organization can appropriately address individual performance while still investigating whether the quality system contributed to the event. Do not substitute one for the other.
A Useful Way to Think About Causation
- Event — what happened? The wrong raw material was added to the process.
- Immediate cause — what directly produced the event? The operator selected Material 4286 instead of Material 4826.
- Contributing conditions — what made the error more likely or harder to detect? Similar packaging, adjacent storage, similar part numbers, manual verification, interruptions, inadequate visual differentiation.
- Systemic / underlying cause — why did the process permit those conditions to exist without adequate prevention or detection? The material-control process relied almost entirely on operator visual identification despite foreseeable confusion between similar materials and lacked an independent or technological verification control.
This final level is where meaningful corrective-action opportunities often become visible.
The “Human Error → Retrain” Loop
- Human makes error
- Deviation opened
- Root cause = human error
- Employee retrained
- Record closed
- Underlying process remains unchanged
- Different employee encounters same conditions
- Problem recurs
Key takeaway
If recurrence can move from one trained employee to another without the system changing, the system deserves investigation.
Root Cause Analysis Should Follow Evidence, Not Intuition
An investigator often hears the description of an event and immediately develops a theory. That is normal. The mistake is turning that theory into a conclusion before gathering evidence.
Start with facts. Ask: What actually happened? What should have happened? What evidence proves each statement? Where did expected performance first deviate? Which controls should have prevented the event? Which controls should have detected it? Did those controls exist? Did they function? Were they realistically designed?
Root cause analysis should be an evidence-building process, not a search for facts that support the first explanation.
FDA has also described root cause analysis as a structured analytical approach intended to understand why an incident occurred so that targeted interventions can help prevent similar events.[2] FDA's current food-safety RCA materials describe practices including defining the problem, gathering relevant information, analyzing data, considering multiple possible causes, using tools such as process maps and cause-and-effect diagrams, identifying root cause or causes, taking action, and verifying that actions worked.
Source: U.S. Food and Drug Administration, Strengthening Food Safety through Root Cause Analysis. This FDA resource concerns food safety and is cited for its general RCA methodology—not as a universal regulatory requirement for all industries.
Step 1: Write a Neutral Problem Statement
Do not hide your presumed root cause inside the problem statement.
Weak
“Operator failed to follow the SOP due to lack of attention.” That statement already assumes the SOP was adequate, the operator's attention was the problem, and no other factors contributed. None of that may have been established yet.
Better
“On August 12, an operator recorded lot 4826 on the batch record although lot 4286 was used during the operation.” Now the investigation has a factual starting point.
A strong problem statement generally establishes what occurred, what should have occurred, when, where, the product/process/system involved, how the issue was detected, the known scope, and the potential impact. Keep causation out of the statement until causation has been investigated.
Step 2: Protect the Process First
Root cause analysis can take time. Risk may need to be controlled immediately. Possible containment or correction activities include stopping affected operations, placing product on hold, segregating materials, correcting a record appropriately, blocking a system configuration, removing defective equipment from service, increasing temporary verification, and notifying affected groups.
But do not confuse containment with corrective action.
| Action | Question it answers |
| Containment | How do we control the immediate risk? |
| Correction | How do we fix this specific occurrence? |
| Investigation | Why did this happen? |
| Corrective action | What should change because of what we learned? |
| Effectiveness check | Did the change actually work? |
Terminology differs between frameworks and organizations. Preserve that nuance.
Step 3: Build the Timeline
Create a chronological sequence of events before debating causes. A useful timeline may include the procedure revision in effect, employee training status, work start, material selection, equipment state, software activity, hand-offs, environmental conditions, interruptions, inspections, alarm events, approvals, deviations from the expected sequence, and detection of the issue.
Timelines are powerful because they expose hidden dependencies. For example, the investigation begins with “Employee skipped the independent verification.” The timeline reveals:
- The normal verifier called out sick.
- No alternate verifier was assigned.
- Production was expected to continue.
- The procedure did not describe what to do when a verifier was unavailable.
- The system allowed the process to proceed without the verification.
The missed check is still real. But the investigation now has substantially more useful information.
Step 4: Look at the Work as It Is Actually Performed
Do not investigate only from behind a desk. A procedure may look perfectly clear to the person reading it in Quality. Go to the point of use.
Where appropriate, observe the task, look at the equipment, review the screen, compare the materials, follow the actual workflow, look at the physical space, evaluate hand-offs, and ask the employee to show how the process works. You may discover that:
- The form does not follow the order of operations.
- Required information is located across the room.
- Two buttons look almost identical.
- A field defaults to the wrong value.
- Operators routinely create unofficial workarounds because the approved process is impractical.
- A verification step occurs after the point where it could actually prevent the failure.
A root cause investigation should study the process that exists in reality—not only the process described on paper.
Step 5: Interview Without Turning the Interview Into an Interrogation
Employee interviews are valuable sources of evidence. But the goal is to understand the event. Start with: “Walk me through what happened.”
Then explore: What did you expect to happen? What information were you using? What was different from normal? Was anything unclear? Were you interrupted? Was there time pressure? Have you encountered this situation before? Is there a workaround employees commonly use? What makes this task difficult? What would make this mistake harder to make?
Avoid beginning with “Why didn't you follow the SOP?” That question assumes the investigation has already established the cause. It can also discourage employees from telling you about process weaknesses that may matter more than the individual mistake.
Step 6: Ask Why the Control System Failed
For every meaningful failure, consider three layers:
Prevention
What should have prevented this error?
Detection
If prevention failed, what should have detected it before impact?
Response
Once detected, what should have limited the consequence?
Example — wrong material selected:
| Layer | Control that should have applied |
| Prevention | Physical segregation, barcode scanning, clear labeling |
| Detection | Independent verification before use |
| Response | In-process testing or reconciliation before release |
If all three layers fail, the investigation should not focus exclusively on the person who made the first mistake. The control architecture itself deserves attention.
10 Questions to Ask Before Calling It “Human Error”
- Was the procedure clear, current, and realistically executable?
- Was the employee appropriately trained and competent for the task?
- Was the correct information available at the point of use?
- Did the process depend unnecessarily on memory or manual transcription?
- Could equipment or software design have prevented or detected the error?
- Were materials, labels, fields, or choices easy to confuse?
- Did workload, interruptions, fatigue, staffing, or competing priorities contribute?
- Was an appropriate independent or automated verification available?
- Has the same or a similar event occurred before—even under a different classification?
- Why did the QMS fail to prevent or detect the mistake before it became a quality event?
Bottom line
If those questions have not been evaluated, “human error” is probably the beginning of the investigation—not the conclusion.
Step 7: Use 5 Whys Correctly
The 5 Whys is useful because it forces the investigator past the immediate event. But it is not magic.
Problem: Wrong label applied.
- Why? Operator selected the wrong label.
- Why? Two similar labels were present.
- Why? Unused labels from the previous product remained at the workstation.
- Why? Line clearance did not require independent verification of label removal.
- Why? The process relied on operator confirmation without a second preventive or detective control.
Now the investigation has found something the organization can change. But sometimes three Whys are sufficient, sometimes more than five are necessary, and sometimes several causal paths exist. The objective is not to reach exactly five boxes. The objective is to follow causality as far as the evidence supports.
Step 8: Use a Fishbone to Expand the Investigation
For a complex event, use a cause-and-effect / Ishikawa approach to explore possible contributors.
| Category | Prompts to explore |
| People | Training, competency, experience, communication, workload, supervision |
| Process | Sequence, procedure clarity, hand-offs, verification, complexity, approval logic |
| Equipment / Technology | Equipment design, user interface, alarms, configuration, maintenance, calibration |
| Materials | Identification, similarity, storage, labeling, supplier variation |
| Measurement | Methods, sampling, acceptance criteria, instrument capability, data interpretation |
| Environment | Lighting, noise, layout, interruptions, temperature, material flow |
| Management / System | Staffing, scheduling, responsibilities, oversight, change control, risk management, resources |
Important
A fishbone generates hypotheses. It does not generate root causes. Each suspected cause still needs evidence.
Step 9: Don't Default to Retraining
Retraining has its place. It is reasonable when the evidence demonstrates a genuine knowledge or competency gap—for example, the requirement was clear, the process was appropriately designed, the employee had not been trained, the employee misunderstood a defined requirement, and the issue was reasonably isolated.
But retraining does not fix confusing software, look-alike materials, an impractical procedure, excessive manual transcription, missing controls, poor workflow design, unclear responsibilities, inadequate equipment, or chronic understaffing.
Ask this
If a perfectly trained new employee started tomorrow, would the process still make the same mistake reasonably easy to make?
If yes, training alone probably does not address the underlying weakness.
Side-by-Side Example: Two Investigations of the Same Event
Investigation A: Stops at the Person
Event: Incorrect component installed.
Root Cause: Operator failed to follow the component-selection procedure.
CAPA: Retrain operator.
Effectiveness Check: Training completed.
What is missing? Why was the component confused? Could another operator make the same mistake? Did the process have preventive controls? Did detection controls function? Has this occurred before? Does training completion demonstrate effectiveness?
Investigation B: Examines the System
Event: Component 4826 installed instead of Component 4286.
Evidence:
- Components stored in adjacent bins
- Visually similar packaging
- Part numbers differ by transposed digits
- Manual visual selection
- No barcode verification
- No independent check before installation
- Prior near-miss involving the same component pair
Immediate cause: Incorrect component selected.
Contributing factors: Similarity, storage location, manual identification, inadequate verification.
Systemic cause: Material-control design relied primarily on manual visual differentiation of known look-alike components without sufficient error-prevention or detection controls.
Potential actions: Segregate look-alike materials, improve labeling, implement barcode verification, add risk-based verification, update material-location standards, review other similar components, and train employees on the revised controls.
Effectiveness measure: Confirm revised controls are consistently used and monitor applicable transactions and near-misses for a period or sample size capable of detecting recurrence.
The difference
Investigation B does not excuse the operator's mistake. It gives the organization something meaningful to improve.
Step 10: Look Horizontally, Not Just at the One Event
Once a cause is identified, ask: Where else does this condition exist? This is one of the differences between correcting an event and improving a quality system.
- Similar material labels → review other look-alike materials.
- Confusing SOP structure → review similar SOPs.
- Software configuration → review other workflows using the same configuration.
- Weak training assignment logic → review other controlled-document training.
- Supplier failure → review similar suppliers or materials.
- Equipment design → review similar equipment.
The original event defines where the investigation started. It should not automatically define where the investigation ends.
Step 11: Search for Recurrence Using More Than the Event Title
Do not assume an event is isolated because no previous record has the same title. Search deviations, nonconformances, CAPAs, complaints, audit findings, OOS/OOT records where applicable, supplier issues, maintenance records, near-misses, and help-desk tickets where relevant.
A single systemic problem may previously have been categorized as operator error, documentation error, a training issue, incorrect selection, missed verification, or procedure deviation. Different labels can hide the same failure mode. Trending should look for patterns in causes and failure mechanisms, not only matching titles.
Step 12: Make the CAPA Logically Follow the Cause
| Cause identified | Action may involve |
| Procedure unclear | Procedure redesign |
| Competency gap | Targeted training + competency verification |
| Similar materials | Segregation / visual differentiation / mistake-proofing |
| Manual transcription | Automation or independent verification |
| Poor interface | Software or equipment redesign |
| Missing control | Preventive or detective control |
| Supplier weakness | Supplier-control improvement |
| Workflow overload | Process or resource redesign |
| Software configuration | Configuration change + appropriate testing |
Test your logic
If you cannot draw a clear line from the cause to the corrective action, reconsider either the cause or the action.
Step 13: Don't Confuse Implementation With Effectiveness
Implementation check
Did we do what we said we would do? Example: all employees completed training.
Effectiveness check
Did the action actually reduce the problem? Example: no recurrence of the targeted error across an appropriate number of opportunities, and the new control was confirmed to be functioning as intended.
Key distinction
Training completion proves training occurred. It does not prove that the original problem was eliminated.
The effectiveness method should consider the frequency of the event, process volume, severity, detectability, the expected recurrence interval, the nature of the control, and the available data.
Do not automatically use an arbitrary 30-day effectiveness period. If the historical failure occurred once every 18 months, a 30-day review may tell you almost nothing.
What If You Cannot Identify a Definitive Root Cause?
Sometimes the evidence does not support certainty. Do not manufacture certainty to complete a CAPA form. A defensible investigation may conclude that the root cause is inconclusive.
In that situation: document what was evaluated, document the available evidence, identify potential causes considered, explain which were ruled out and why, identify remaining plausible causes, evaluate risk, implement appropriate controls where warranted, increase monitoring if justified, trend for recurrence, and reopen or expand the investigation if new evidence appears.
A well-supported inconclusive finding is better than an unsupported statement presented as fact. This is consistent with FDA's approach in the cited Pharmathen warning letter, where FDA specifically expected broader manufacturing review when a laboratory cause could not be conclusively established. Do not overgeneralize that pharmaceutical-specific expectation to every QMS framework.
When Can Human Error Actually Be the Finding?
There are cases where evidence may support an individual behavioral cause. For example, the investigation may establish that:
- The process was appropriately designed.
- The instructions were clear.
- The employee was trained and competent.
- Required tools and information were available.
- Controls were functioning.
- Workload and environmental conditions were reasonable.
- The employee knowingly departed from established requirements.
- No relevant systemic contributor was identified.
That can be a legitimate finding. The point is not “human error can never be the cause.” The point is that human error should be demonstrated through investigation, not selected as the default explanation because a person happened to be involved.
Common RCA Failures
| Failure | What it usually means |
| Root cause = human error | Investigation may have stopped at the immediate action |
| Root cause = SOP not followed | Describes nonconformance, not necessarily cause |
| CAPA = retraining | May not address the identified system weakness |
| 5 Whys completed mechanically | Form completed without causal analysis |
| Only employee interviews used | Objective system evidence may be missing |
| No historical review | Recurrence may go undetected |
| No horizontal review | Same weakness may remain elsewhere |
| Effectiveness = action complete | Implementation confused with effectiveness |
| Predetermined conclusion | Investigation becomes confirmation bias |
| Endless analysis | Investigation effort exceeds what risk justifies |
A Stronger RCA Model
- 1. Define the event objectively
- 2. Contain immediate risk
- 3. Preserve and gather evidence
- 4. Build the timeline
- 5. Observe the actual process
- 6. Identify prevention and detection controls
- 7. Develop causal hypotheses
- 8. Test hypotheses against evidence
- 9. Identify immediate, contributing, and systemic causes
- 10. Evaluate recurrence and horizontal scope
- 11. Select risk-proportionate actions
- 12. Implement
- 13. Verify effectiveness
- 14. Trend for recurrence
The Question That Changes the Investigation
Don't stop with
“Why did the employee make the mistake?”
Ask
“Why did our process allow one human mistake to become a quality failure?”
People will make mistakes. An effective QMS does not assume otherwise. It establishes processes that make important errors less likely, makes failures more detectable, limits their consequences, and learns when the controls are not enough.
That is why root cause analysis should not be an exercise in assigning blame. It should be an exercise in learning how the failure happened well enough to prevent it from happening again.
And if your CAPA repeatedly ends with Human error → Retrain → Close while similar problems continue to recur, the problem may no longer be the individual investigation. It may be the CAPA system itself.
FAQ
Is human error a valid root cause?
Sometimes, but it should be supported by evidence rather than used as a default classification. Even when an individual's action directly caused an event, the investigation should consider whether process, system, procedural, equipment, environmental, or organizational factors contributed.
Why isn't retraining always an effective CAPA?
Retraining addresses knowledge or competency. It does not redesign confusing processes, eliminate look-alike materials, change poor interfaces, automate manual transcription, or add missing controls.
Is 5 Whys required for root cause analysis?
No. It is one tool. Depending on the event, organizations may use timelines, process mapping, cause-and-effect diagrams, fault-tree analysis, barrier analysis, or other methods.
What is the difference between an immediate cause and root cause?
An immediate cause is the condition or action directly associated with the event. Root or systemic causal analysis attempts to determine why that condition existed and what underlying weakness allowed it to result in the failure.
Can a root cause investigation be inconclusive?
Yes. When the available evidence does not support a definitive cause, documenting uncertainty and the evidence evaluated is preferable to inventing an unsupported conclusion.
How should CAPA effectiveness be verified?
The effectiveness check should evaluate whether the action addressed the identified cause and achieved the intended result—not merely whether the action was completed.
How Athyrion Can Help
Recurring deviations, repeated “human error” investigations, and CAPAs that close without demonstrating effectiveness can indicate a deeper quality-system problem. Athyrion helps organizations strengthen investigations, root cause analysis, CAPA, quality procedures, risk-based quality processes, QMS implementation, and digital quality workflows.
Athyrion does not guarantee regulatory, certification, or audit outcomes.
References
- U.S. Food and Drug Administration. Pharmathen International S.A. — Warning Letter 320-26-80. May 27, 2026.
- U.S. Food and Drug Administration. Strengthening Food Safety through Root Cause Analysis.