We’ve all experienced the “Groundhog Day” feeling on the plant floor. A critical centrifugal pump goes down, the maintenance team swaps the bearing, realigns the shaft, and puts it back in service. Three months later, the exact same bearing seizes.
When this happens, it’s easy to blame bad parts or bad luck. But in reliability engineering, a repeat failure is rarely a mechanical coincidence. It is hard evidence that your previous investigation stopped at the physical root cause and completely missed the latent root cause.
The physical root cause tells you what broke (e.g., the bearing seized due to lack of lubrication). The latent root cause tells you why the system allowed it to break (e.g., the lubricant drum was stored outside without a desiccant breather, allowing moisture ingress, because the procurement spec didn’t require one).
If you only fix the physical root cause, you’re just applying a band-aid. To actually eliminate the defect, you have to investigate the system. Here is a practical, technical framework for how plants should investigate repeat equipment failures, moving past the basics and into real defect elimination.
Securing the “Crime Scene” Before the Scrap Bin Claims It
The biggest mistake in a failure investigation happens in the first thirty minutes after a breakdown. It’s 2:00 AM, production is halted, and the pressure to get the line running is immense. The technicians strip the failed assembly, toss the damaged parts into the scrap bin, and install the spares.
By the time the reliability engineer walks onto the floor the next morning, the forensic evidence is gone. You cannot investigate a failure if you throw away the evidence.
Physical Preservation
Before any corrective repair begins, you need a strict “bag and tag” protocol for failed components. If a bearing fails, do not let the technician wash it in solvent to “see it better.” Wash away the grease, and you wash away the wear particles that tell the story of the failure. Place the component in a sealed, static-free bag. Tag it with the asset ID, the date, the operating hours, and the name of the technician who pulled it.
If the failure involves lubrication, pull tribology samples from the sump and the distribution lines before the system is flushed. If it’s a catastrophic failure like a sheared shaft or a cracked housing, preserve the fracture faces. You may need metallurgical analysis or Scanning Electron Microscopy (SEM) later, and a wire brush applied by a well-meaning mechanic will destroy the fracture morphology.
Digital Preservation
Physical evidence tells you what broke; digital evidence tells you how the degradation started. Modern plants generate massive amounts of data, but it’s usually on a rolling overwrite cycle.
When a bad actor fails, immediately freeze the SCADA or DCS historian data for the 72 hours preceding the event. Look for the subtle precursors: a slow drift in motor current signature, a slight increase in discharge pressure, or a pattern of low-priority alarms that operators kept acknowledging and clearing. Export the control system alarm logs. Often, the machine was screaming for help for weeks, but the signals were buried in the alarm flood.
Moving Beyond the “5 Whys” for Complex Failures
Once the evidence is secured, the team needs to sit down and figure out what happened. Many plants default to the “5 Whys” technique. While 5 Whys is great for simple, linear problems (like a blown fuse), it is fundamentally inadequate for complex, repeat equipment failures. It forces you down a single, linear causal path and almost always ends with a dead-end like “operator error” or “technician didn’t follow the SOP.”
For repeat failures, you need a more robust analytical engine.
Timeline Reconstruction
Start by building a chronological map. Don’t just look at the machine; look at the plant. Map operational events against maintenance history. Did the raw material supplier change two weeks ago? Was there a process upset that caused a sudden load spike? Did a new technician perform the last PM, and was there a “No Fault Found” note on a recent work order? Hidden correlations live in the timeline.
Fault Tree Analysis (FTA) and Cause Mapping
Instead of asking “Why?” five times, use logic gates. Fault Tree Analysis forces you to map out all the conditions that must exist simultaneously for a failure to occur using AND/OR gates.
For example, for a pump to run dry and destroy its seal, two things must happen: the flow must stop (Condition A) AND the low-flow interlock must fail to trip the motor (Condition B). By mapping it this way, you stop looking for a single “silver bullet” root cause and start looking at the systemic vulnerabilities that allowed multiple safeguards to fail at once.
Interrogating Human Factors (Stop Blaming the Operator)
If your Root Cause Analysis (RCA) concludes with “operator error,” your investigation is incomplete. Human error is a symptom, not a root cause.
If three different operators make the exact same “mistake,” you don’t have a people problem; you have a design problem. Ask the hard questions: Was the Human-Machine Interface (HMI) confusing? Was the maintenance tooling inadequate, forcing the mechanic to improvise? Was the SOP ambiguous? Was the technician working a double shift? You have to investigate the context of the error to engineer it out of the system.
Closing the Loop: Verification and Weibull Analysis
Here is where most RCAs die. The team identifies a great latent root cause, they write a corrective action plan (CAPA), they update a procedure, and they close the ticket in the CMMS.
But a corrective action is just a hypothesis. How do you mathematically prove that your fix actually worked?
The Verification Gap
Closing a CAPA ticket doesn’t mean the problem is solved; it just means the paperwork is done. When an asset undergoes a major corrective overhaul based on an RCA, it needs to go on a “Watch List.” Mandate a 90- to 180-day intensified condition monitoring protocol. Increase the vibration routes, do weekly oil sampling, or deploy temporary temperature sensors. You are looking for the specific degradation signature that caused the original failure to ensure it hasn’t returned.
Tracking the Weibull Shape Parameter (β)
For critical assets with a nasty habit of failing, reliability engineers should use Weibull analysis to track the failure distribution before and after the fix.
In Weibull analysis, the shape parameter (β) tells you the failure pattern. If β is less than 1, you have infant mortality (usually caused by bad installation or defective parts). If β is greater than 1, you have a wear-out pattern.
When you implement a systemic fix, you should see the Weibull plot shift. If you fix a contamination issue, the early-life failures (β<1) should drop off. If your corrective action is successful, the asset’s failure distribution should stabilize. If you check the data six months later and the β parameter hasn’t moved, your root cause was wrong, or your corrective action wasn’t implemented properly.
Architecting the Digital Backbone for RCA
You cannot execute this level of forensic rigor if your data is trapped in silos. You can’t do a proper timeline reconstruction if your CMMS work order history just says “Replaced Pump,” your SCADA data is locked in a vendor portal, and your vibration data lives on a contractor’s laptop.
To investigate repeat failures effectively, you need a unified digital backbone. This is where a platform like TeroTAM shifts from being a simple work-order tool to a critical reliability engine.
- Unified Asset Genealogy: TeroTAM links the physical work order history, the condition monitoring data, and the operator logs into a single asset timeline. When a bad actor fails, the reliability engineer doesn’t spend two days hunting down spreadsheets; the complete forensic record is pulled up in one view.
- Automated Bad Actor Flagging: You shouldn’t have to manually hunt for repeat failures. TeroTAM enforces structured failure coding at work order closure. When an asset accumulates multiple failures within a rolling window, the platform automatically flags it as a “Bad Actor” and triggers a mandatory RCA workflow, ensuring the issue isn’t buried under routine corrective work.
- CAPA Tracking and Verification: Once a corrective action is logged, TeroTAM automatically generates the follow-up condition monitoring tasks at 30, 60, and 90 days. It tracks the asset’s performance against the expected recovery baseline, ensuring the “Watch List” is actually enforced.
Without this digital architecture, RCA becomes a sporadic, heroic effort that relies on one reliability engineer’s memory and a messy Excel tracker. With it, defect elimination becomes a repeatable, scalable process.
Shifting the Mindset from MTTR to MTBF
A repeat failure is an expensive tuition fee. The machine broke, production stopped, and the plant lost money. The only way to get a return on that painful investment is to extract every ounce of data from the event and permanently engineer the defect out of the facility.
Plants that treat repeat failures as routine maintenance events will always be trapped in a reactive cycle, measuring their success by Mean Time To Repair (MTTR)—how fast they can fix what keeps breaking.
Plants that adopt a rigorous, forensic investigation framework measure their success by Mean Time Between Failures (MTBF). They understand that the goal isn’t to get better at fixing broken machines; the goal is to build a system where the machines don’t break in the first place.
Summing it up
Stop treating the symptoms and start eliminating the defects. Schedule a quick demo now and Discover how TeroTAM’s unified asset history, automated Bad Actor flagging, and CAPA verification workflows provide the digital backbone for world-class Root Cause Analysis.