CAPA & Investigations

How to Perform an Effective Root Cause Analysis Without Blaming “Human Error”

“Human error” may explain what happened, but it often fails to explain why. Learn how to investigate the process, system, equipment, training, and organizational conditions behind quality failures.

An employee selects the wrong component. A technician records the wrong value. An operator skips a step. A reviewer misses an error. A required check is not completed.

The investigation opens. A few days later, the conclusion reads:

Root Cause: Human Error. Corrective Action: Retrain the employee. CAPA closed.

And six months later, something very similar happens again.

That pattern is one of the clearest signs that an investigation stopped at the event instead of finding the conditions that allowed the event to happen.

“Human error” can accurately describe what an individual did. But it often does not explain why the quality system allowed one ordinary mistake to become a product, process, data, or compliance failure.

A stronger investigation asks: Why was this mistake possible? Why wasn't it prevented? Why wasn't it detected earlier? Could the same conditions affect someone else tomorrow?

Those questions are where root cause analysis begins.


The Person Closest to the Failure Is Not Automatically the Root Cause

Start here

The person who made the final mistake may be the last link in the causal chain, not the first.

Consider a wrong-component event. The employee physically selected the wrong item. That is factual. But the investigation may also find that:

  • Two components had nearly identical packaging.
  • Part numbers differed by only two digits.
  • The materials were stored next to each other.
  • The procedure relied entirely on visual verification.
  • There was no barcode or system verification.
  • A previous near-miss involving the same components had already occurred.
  • Production schedules encouraged rapid changeover.
  • No independent check existed before use.

Saying “the operator selected the wrong part” identifies what happened. It does not explain why the process was designed so that one selection mistake could pass through every available control.

That distinction is the heart of effective root cause analysis.


FDA Has Made This Point Explicitly

A 2026 FDA Warning Letter Provides a Clear Example

In a May 27, 2026 warning letter to pharmaceutical manufacturer Pharmathen International S.A., FDA addressed recurring sterility and media-fill failures in which operator error had contributed significantly.[1]

FDA did not simply accept “operator error” as the end of the investigation. The agency stated that the firm's investigation:

FDA / Regulatory Example

“did not include in-depth analysis of the conditions that create the circumstances for human error.”

— U.S. Food and Drug Administration, Warning Letter to Pharmathen International S.A., May 27, 2026

FDA went on to state that an in-depth root-cause analysis was needed to identify effective CAPA, including consideration of design remediation. The agency also criticized the firm's proposed CAPAs because they lacked sufficient detail about how effectiveness would be evaluated.

This is an important distinction. The regulatory concern was not merely that employees had made errors. The concern was that the investigation had not adequately examined why those errors were recurring or what conditions within the operation made them possible.

Source: U.S. Food and Drug Administration, Pharmathen International S.A. — Warning Letter 320-26-80, May 27, 2026.

Scope note

This warning letter concerns pharmaceutical CGMP requirements under 21 CFR Parts 210 and 211. It should not be read to mean that every statement in the letter establishes a universal requirement for ISO 9001, ISO/IEC 17025, medical-device, laboratory, or other QMS environments. It is cited here as a current FDA enforcement example demonstrating why investigations may need to look beyond the immediate human action.


“Human Error” Is Often a Stopping Point Disguised as a Cause

When an investigation concludes “human error,” ask what would happen if the employee were replaced tomorrow. Would the next appropriately trained employee face:

  • The same confusing screen?
  • The same similar labels?
  • The same unclear instruction?
  • The same manual transcription?
  • The same workload?
  • The same missing verification?
  • The same equipment limitation?
  • The same opportunity to make the same mistake?

If the answer is yes, then replacing or retraining the employee has not removed the condition that produced the failure.

That does not mean the employee bears no responsibility. It means accountability and causality are different questions. An organization can appropriately address individual performance while still investigating whether the quality system contributed to the event. Do not substitute one for the other.


A Useful Way to Think About Causation

  1. Event — what happened? The wrong raw material was added to the process.
  2. Immediate cause — what directly produced the event? The operator selected Material 4286 instead of Material 4826.
  3. Contributing conditions — what made the error more likely or harder to detect? Similar packaging, adjacent storage, similar part numbers, manual verification, interruptions, inadequate visual differentiation.
  4. Systemic / underlying cause — why did the process permit those conditions to exist without adequate prevention or detection? The material-control process relied almost entirely on operator visual identification despite foreseeable confusion between similar materials and lacked an independent or technological verification control.

This final level is where meaningful corrective-action opportunities often become visible.


The “Human Error → Retrain” Loop

  • Human makes error
  • Deviation opened
  • Root cause = human error
  • Employee retrained
  • Record closed
  • Underlying process remains unchanged
  • Different employee encounters same conditions
  • Problem recurs
Key takeaway

If recurrence can move from one trained employee to another without the system changing, the system deserves investigation.


Root Cause Analysis Should Follow Evidence, Not Intuition

An investigator often hears the description of an event and immediately develops a theory. That is normal. The mistake is turning that theory into a conclusion before gathering evidence.

Start with facts. Ask: What actually happened? What should have happened? What evidence proves each statement? Where did expected performance first deviate? Which controls should have prevented the event? Which controls should have detected it? Did those controls exist? Did they function? Were they realistically designed?

Root cause analysis should be an evidence-building process, not a search for facts that support the first explanation.

FDA has also described root cause analysis as a structured analytical approach intended to understand why an incident occurred so that targeted interventions can help prevent similar events.[2] FDA's current food-safety RCA materials describe practices including defining the problem, gathering relevant information, analyzing data, considering multiple possible causes, using tools such as process maps and cause-and-effect diagrams, identifying root cause or causes, taking action, and verifying that actions worked.

Source: U.S. Food and Drug Administration, Strengthening Food Safety through Root Cause Analysis. This FDA resource concerns food safety and is cited for its general RCA methodology—not as a universal regulatory requirement for all industries.


Step 1: Write a Neutral Problem Statement

Do not hide your presumed root cause inside the problem statement.

Weak

“Operator failed to follow the SOP due to lack of attention.” That statement already assumes the SOP was adequate, the operator's attention was the problem, and no other factors contributed. None of that may have been established yet.

Better

“On August 12, an operator recorded lot 4826 on the batch record although lot 4286 was used during the operation.” Now the investigation has a factual starting point.

A strong problem statement generally establishes what occurred, what should have occurred, when, where, the product/process/system involved, how the issue was detected, the known scope, and the potential impact. Keep causation out of the statement until causation has been investigated.


Step 2: Protect the Process First

Root cause analysis can take time. Risk may need to be controlled immediately. Possible containment or correction activities include stopping affected operations, placing product on hold, segregating materials, correcting a record appropriately, blocking a system configuration, removing defective equipment from service, increasing temporary verification, and notifying affected groups.

But do not confuse containment with corrective action.

ActionQuestion it answers
ContainmentHow do we control the immediate risk?
CorrectionHow do we fix this specific occurrence?
InvestigationWhy did this happen?
Corrective actionWhat should change because of what we learned?
Effectiveness checkDid the change actually work?

Terminology differs between frameworks and organizations. Preserve that nuance.


Step 3: Build the Timeline

Create a chronological sequence of events before debating causes. A useful timeline may include the procedure revision in effect, employee training status, work start, material selection, equipment state, software activity, hand-offs, environmental conditions, interruptions, inspections, alarm events, approvals, deviations from the expected sequence, and detection of the issue.

Timelines are powerful because they expose hidden dependencies. For example, the investigation begins with “Employee skipped the independent verification.” The timeline reveals:

  • The normal verifier called out sick.
  • No alternate verifier was assigned.
  • Production was expected to continue.
  • The procedure did not describe what to do when a verifier was unavailable.
  • The system allowed the process to proceed without the verification.

The missed check is still real. But the investigation now has substantially more useful information.


Step 4: Look at the Work as It Is Actually Performed

Do not investigate only from behind a desk. A procedure may look perfectly clear to the person reading it in Quality. Go to the point of use.

Where appropriate, observe the task, look at the equipment, review the screen, compare the materials, follow the actual workflow, look at the physical space, evaluate hand-offs, and ask the employee to show how the process works. You may discover that:

  • The form does not follow the order of operations.
  • Required information is located across the room.
  • Two buttons look almost identical.
  • A field defaults to the wrong value.
  • Operators routinely create unofficial workarounds because the approved process is impractical.
  • A verification step occurs after the point where it could actually prevent the failure.

A root cause investigation should study the process that exists in reality—not only the process described on paper.


Step 5: Interview Without Turning the Interview Into an Interrogation

Employee interviews are valuable sources of evidence. But the goal is to understand the event. Start with: “Walk me through what happened.”

Then explore: What did you expect to happen? What information were you using? What was different from normal? Was anything unclear? Were you interrupted? Was there time pressure? Have you encountered this situation before? Is there a workaround employees commonly use? What makes this task difficult? What would make this mistake harder to make?

Avoid beginning with “Why didn't you follow the SOP?” That question assumes the investigation has already established the cause. It can also discourage employees from telling you about process weaknesses that may matter more than the individual mistake.


Step 6: Ask Why the Control System Failed

For every meaningful failure, consider three layers:

Prevention

What should have prevented this error?

Detection

If prevention failed, what should have detected it before impact?

Response

Once detected, what should have limited the consequence?

Example — wrong material selected:

LayerControl that should have applied
PreventionPhysical segregation, barcode scanning, clear labeling
DetectionIndependent verification before use
ResponseIn-process testing or reconciliation before release

If all three layers fail, the investigation should not focus exclusively on the person who made the first mistake. The control architecture itself deserves attention.


10 Questions to Ask Before Calling It “Human Error”

  1. Was the procedure clear, current, and realistically executable?
  2. Was the employee appropriately trained and competent for the task?
  3. Was the correct information available at the point of use?
  4. Did the process depend unnecessarily on memory or manual transcription?
  5. Could equipment or software design have prevented or detected the error?
  6. Were materials, labels, fields, or choices easy to confuse?
  7. Did workload, interruptions, fatigue, staffing, or competing priorities contribute?
  8. Was an appropriate independent or automated verification available?
  9. Has the same or a similar event occurred before—even under a different classification?
  10. Why did the QMS fail to prevent or detect the mistake before it became a quality event?
Bottom line

If those questions have not been evaluated, “human error” is probably the beginning of the investigation—not the conclusion.


Step 7: Use 5 Whys Correctly

The 5 Whys is useful because it forces the investigator past the immediate event. But it is not magic.

Problem: Wrong label applied.

  • Why? Operator selected the wrong label.
  • Why? Two similar labels were present.
  • Why? Unused labels from the previous product remained at the workstation.
  • Why? Line clearance did not require independent verification of label removal.
  • Why? The process relied on operator confirmation without a second preventive or detective control.

Now the investigation has found something the organization can change. But sometimes three Whys are sufficient, sometimes more than five are necessary, and sometimes several causal paths exist. The objective is not to reach exactly five boxes. The objective is to follow causality as far as the evidence supports.


Step 8: Use a Fishbone to Expand the Investigation

For a complex event, use a cause-and-effect / Ishikawa approach to explore possible contributors.

CategoryPrompts to explore
PeopleTraining, competency, experience, communication, workload, supervision
ProcessSequence, procedure clarity, hand-offs, verification, complexity, approval logic
Equipment / TechnologyEquipment design, user interface, alarms, configuration, maintenance, calibration
MaterialsIdentification, similarity, storage, labeling, supplier variation
MeasurementMethods, sampling, acceptance criteria, instrument capability, data interpretation
EnvironmentLighting, noise, layout, interruptions, temperature, material flow
Management / SystemStaffing, scheduling, responsibilities, oversight, change control, risk management, resources
Important

A fishbone generates hypotheses. It does not generate root causes. Each suspected cause still needs evidence.


Step 9: Don't Default to Retraining

Retraining has its place. It is reasonable when the evidence demonstrates a genuine knowledge or competency gap—for example, the requirement was clear, the process was appropriately designed, the employee had not been trained, the employee misunderstood a defined requirement, and the issue was reasonably isolated.

But retraining does not fix confusing software, look-alike materials, an impractical procedure, excessive manual transcription, missing controls, poor workflow design, unclear responsibilities, inadequate equipment, or chronic understaffing.

Ask this

If a perfectly trained new employee started tomorrow, would the process still make the same mistake reasonably easy to make?

If yes, training alone probably does not address the underlying weakness.


Side-by-Side Example: Two Investigations of the Same Event

Investigation A: Stops at the Person

Event: Incorrect component installed.
Root Cause: Operator failed to follow the component-selection procedure.
CAPA: Retrain operator.
Effectiveness Check: Training completed.

What is missing? Why was the component confused? Could another operator make the same mistake? Did the process have preventive controls? Did detection controls function? Has this occurred before? Does training completion demonstrate effectiveness?

Investigation B: Examines the System

Event: Component 4826 installed instead of Component 4286.

Evidence:

  • Components stored in adjacent bins
  • Visually similar packaging
  • Part numbers differ by transposed digits
  • Manual visual selection
  • No barcode verification
  • No independent check before installation
  • Prior near-miss involving the same component pair

Immediate cause: Incorrect component selected.
Contributing factors: Similarity, storage location, manual identification, inadequate verification.
Systemic cause: Material-control design relied primarily on manual visual differentiation of known look-alike components without sufficient error-prevention or detection controls.

Potential actions: Segregate look-alike materials, improve labeling, implement barcode verification, add risk-based verification, update material-location standards, review other similar components, and train employees on the revised controls.

Effectiveness measure: Confirm revised controls are consistently used and monitor applicable transactions and near-misses for a period or sample size capable of detecting recurrence.

The difference

Investigation B does not excuse the operator's mistake. It gives the organization something meaningful to improve.


Step 10: Look Horizontally, Not Just at the One Event

Once a cause is identified, ask: Where else does this condition exist? This is one of the differences between correcting an event and improving a quality system.

  • Similar material labels → review other look-alike materials.
  • Confusing SOP structure → review similar SOPs.
  • Software configuration → review other workflows using the same configuration.
  • Weak training assignment logic → review other controlled-document training.
  • Supplier failure → review similar suppliers or materials.
  • Equipment design → review similar equipment.

The original event defines where the investigation started. It should not automatically define where the investigation ends.


Step 11: Search for Recurrence Using More Than the Event Title

Do not assume an event is isolated because no previous record has the same title. Search deviations, nonconformances, CAPAs, complaints, audit findings, OOS/OOT records where applicable, supplier issues, maintenance records, near-misses, and help-desk tickets where relevant.

A single systemic problem may previously have been categorized as operator error, documentation error, a training issue, incorrect selection, missed verification, or procedure deviation. Different labels can hide the same failure mode. Trending should look for patterns in causes and failure mechanisms, not only matching titles.


Step 12: Make the CAPA Logically Follow the Cause

Cause identifiedAction may involve
Procedure unclearProcedure redesign
Competency gapTargeted training + competency verification
Similar materialsSegregation / visual differentiation / mistake-proofing
Manual transcriptionAutomation or independent verification
Poor interfaceSoftware or equipment redesign
Missing controlPreventive or detective control
Supplier weaknessSupplier-control improvement
Workflow overloadProcess or resource redesign
Software configurationConfiguration change + appropriate testing
Test your logic

If you cannot draw a clear line from the cause to the corrective action, reconsider either the cause or the action.


Step 13: Don't Confuse Implementation With Effectiveness

Implementation check

Did we do what we said we would do? Example: all employees completed training.

Effectiveness check

Did the action actually reduce the problem? Example: no recurrence of the targeted error across an appropriate number of opportunities, and the new control was confirmed to be functioning as intended.

Key distinction

Training completion proves training occurred. It does not prove that the original problem was eliminated.

The effectiveness method should consider the frequency of the event, process volume, severity, detectability, the expected recurrence interval, the nature of the control, and the available data.

Do not automatically use an arbitrary 30-day effectiveness period. If the historical failure occurred once every 18 months, a 30-day review may tell you almost nothing.


What If You Cannot Identify a Definitive Root Cause?

Sometimes the evidence does not support certainty. Do not manufacture certainty to complete a CAPA form. A defensible investigation may conclude that the root cause is inconclusive.

In that situation: document what was evaluated, document the available evidence, identify potential causes considered, explain which were ruled out and why, identify remaining plausible causes, evaluate risk, implement appropriate controls where warranted, increase monitoring if justified, trend for recurrence, and reopen or expand the investigation if new evidence appears.

A well-supported inconclusive finding is better than an unsupported statement presented as fact. This is consistent with FDA's approach in the cited Pharmathen warning letter, where FDA specifically expected broader manufacturing review when a laboratory cause could not be conclusively established. Do not overgeneralize that pharmaceutical-specific expectation to every QMS framework.


When Can Human Error Actually Be the Finding?

There are cases where evidence may support an individual behavioral cause. For example, the investigation may establish that:

  • The process was appropriately designed.
  • The instructions were clear.
  • The employee was trained and competent.
  • Required tools and information were available.
  • Controls were functioning.
  • Workload and environmental conditions were reasonable.
  • The employee knowingly departed from established requirements.
  • No relevant systemic contributor was identified.

That can be a legitimate finding. The point is not “human error can never be the cause.” The point is that human error should be demonstrated through investigation, not selected as the default explanation because a person happened to be involved.


Common RCA Failures

FailureWhat it usually means
Root cause = human errorInvestigation may have stopped at the immediate action
Root cause = SOP not followedDescribes nonconformance, not necessarily cause
CAPA = retrainingMay not address the identified system weakness
5 Whys completed mechanicallyForm completed without causal analysis
Only employee interviews usedObjective system evidence may be missing
No historical reviewRecurrence may go undetected
No horizontal reviewSame weakness may remain elsewhere
Effectiveness = action completeImplementation confused with effectiveness
Predetermined conclusionInvestigation becomes confirmation bias
Endless analysisInvestigation effort exceeds what risk justifies

A Stronger RCA Model

  • 1. Define the event objectively
  • 2. Contain immediate risk
  • 3. Preserve and gather evidence
  • 4. Build the timeline
  • 5. Observe the actual process
  • 6. Identify prevention and detection controls
  • 7. Develop causal hypotheses
  • 8. Test hypotheses against evidence
  • 9. Identify immediate, contributing, and systemic causes
  • 10. Evaluate recurrence and horizontal scope
  • 11. Select risk-proportionate actions
  • 12. Implement
  • 13. Verify effectiveness
  • 14. Trend for recurrence

The Question That Changes the Investigation

Don't stop with

“Why did the employee make the mistake?”

Ask

“Why did our process allow one human mistake to become a quality failure?”

People will make mistakes. An effective QMS does not assume otherwise. It establishes processes that make important errors less likely, makes failures more detectable, limits their consequences, and learns when the controls are not enough.

That is why root cause analysis should not be an exercise in assigning blame. It should be an exercise in learning how the failure happened well enough to prevent it from happening again.

And if your CAPA repeatedly ends with Human error → Retrain → Close while similar problems continue to recur, the problem may no longer be the individual investigation. It may be the CAPA system itself.


FAQ

Is human error a valid root cause?

Sometimes, but it should be supported by evidence rather than used as a default classification. Even when an individual's action directly caused an event, the investigation should consider whether process, system, procedural, equipment, environmental, or organizational factors contributed.

Why isn't retraining always an effective CAPA?

Retraining addresses knowledge or competency. It does not redesign confusing processes, eliminate look-alike materials, change poor interfaces, automate manual transcription, or add missing controls.

Is 5 Whys required for root cause analysis?

No. It is one tool. Depending on the event, organizations may use timelines, process mapping, cause-and-effect diagrams, fault-tree analysis, barrier analysis, or other methods.

What is the difference between an immediate cause and root cause?

An immediate cause is the condition or action directly associated with the event. Root or systemic causal analysis attempts to determine why that condition existed and what underlying weakness allowed it to result in the failure.

Can a root cause investigation be inconclusive?

Yes. When the available evidence does not support a definitive cause, documenting uncertainty and the evidence evaluated is preferable to inventing an unsupported conclusion.

How should CAPA effectiveness be verified?

The effectiveness check should evaluate whether the action addressed the identified cause and achieved the intended result—not merely whether the action was completed.


How Athyrion Can Help

Recurring deviations, repeated “human error” investigations, and CAPAs that close without demonstrating effectiveness can indicate a deeper quality-system problem. Athyrion helps organizations strengthen investigations, root cause analysis, CAPA, quality procedures, risk-based quality processes, QMS implementation, and digital quality workflows.

Athyrion does not guarantee regulatory, certification, or audit outcomes.


References

  1. U.S. Food and Drug Administration. Pharmathen International S.A. — Warning Letter 320-26-80. May 27, 2026.
  2. U.S. Food and Drug Administration. Strengthening Food Safety through Root Cause Analysis.

Are the same problemsshowing up again?

Recurring deviations, repeated “human error” investigations, and CAPAs that close without demonstrating effectiveness can indicate a deeper quality-system problem. Take Athyrion's free 10-question QMS Health Assessment to evaluate key areas of your QMS and identify potential gaps, risks, and improvement priorities.