Picture this. It’s 2:47 a.m. on a Tuesday. A critical centrifugal pump on Line 3 trips again — the fourth time in eighteen months. The maintenance team rolls in, swaps the mechanical seal, checks alignment, logs the work order, and closes it out. The pump hums back to life. Everyone exhales.
And then, three months later, it happens again.
The maintenance manager pulls up the CMMS records. The pump is on a monthly preventive maintenance schedule. Bearings greased. Coupling inspected. Vibration checked. Every box ticked, every month, without fail. So why does the same component keep failing?
This scenario plays out thousands of times a day across manufacturing plants, processing facilities, power stations, and commercial buildings worldwide. Organizations invest heavily in preventive maintenance programs — scheduling software, spare parts inventories, dedicated PM crews — and yet the same assets keep breaking in the same ways.
Industry research paints a sobering picture. Studies in reliability engineering suggest that only 20 to 30 percent of traditional time-based PM tasks actually prevent functional failures. The rest are either performed too early (replacing components with useful life remaining), too late (after degradation has already progressed), or are simply the wrong tasks for the actual failure mode.
So the honest question maintenance leaders must confront is this: If we’re doing preventive maintenance, why does the same equipment keep breaking?
The answer is uncomfortable but important. Preventive maintenance is necessary. But it is structurally insufficient on its own. Repeat failures persist because of deep, systemic gaps in how PM programs are designed, executed, measured, and improved.
Let’s unpack those gaps.
PM Was Never Designed to Be a Standalone Fix
To understand why a PM falls short, it helps to understand where it came from.
Preventive maintenance, as most organizations practice it today, evolved from the calendar-based “lube, inspect, and adjust” routines of the 1950s through the 1970s. The underlying assumption was straightforward: equipment wears out predictably over time, so if you service it at fixed intervals, you prevent failure.
That assumption held reasonably well for simple, mechanically dominated systems. But modern industrial assets are different. They involve complex interactions between mechanical, electrical, electronic, and software subsystems. Their failure modes are influenced by load profiles, ambient conditions, operator behavior, installation quality, and a dozen other variables that a calendar simply cannot account for.
Consider the P-F Curve, a foundational concept in reliability engineering. It describes the interval between the point where a potential failure (P) becomes detectable — say, a slight increase in vibration — and the point where the asset reaches functional failure (F). For some components, that window is weeks. For others, it’s hours. A monthly PM task might land well after the P-point has passed, meaning the technician is inspecting an asset that has already crossed into active degradation. Or it might land so early that the inspection reveals nothing, giving false confidence.
PM, in its traditional form, is a blunt instrument pointed at a problem that often requires a scalpel.
Seven Reasons Repeat Failures Persist Despite PM
1. PM Tasks Address Symptoms, Not Root Causes
This is the most common and most costly gap. A technician replaces a worn bearing during a PM visit. The work order is closed. But nobody asks why the bearing wore out in six months instead of the expected three years. Was it shaft misalignment? Contamination from a damaged seal? Chronic overloading due to a process change? An incorrect lubricant type?
Without a formal Root Cause Analysis (RCA) — a structured 5-Why, a Fishbone diagram, or a failure mode investigation — the underlying cause remains unaddressed. The new bearing is now on the same accelerated path to failure as the old one. The PM schedule keeps the appointment. The failure keeps the appointment too.
2. PM Schedules Are Static and Generic
Most PM programs are built from OEM manuals at the time of commissioning and then left largely unchanged. A motor running 24/7 in a humid, dust-laden cement plant receives the same quarterly inspection frequency as an identical motor running eight hours a day in a climate-controlled pharmaceutical cleanroom.
OEM recommendations are conservative baselines designed for the widest possible range of operating conditions. They do not account for your specific duty cycle, your environment, your load variations. When PM intervals are never recalibrated against actual operating data, they become rituals rather than reliability tools.
3. Poor or Missing Failure History Data
Open the work order history for a frequently failing asset in many organizations, and you’ll find entries like:
- “Fixed.”
- “Replaced part.”
- “Pump repaired, tested OK.”
There is no standardized failure code. No notation of which component failed, what the contributing factor was, whether the repair was a temporary patch or a permanent fix. Without structured, coded failure data, patterns remain invisible. The same failure mode keeps “surprising” the team because the system has no memory. You cannot fix what you cannot see.
4. Compliance ≠ Effectiveness
Walk into a maintenance planning meeting and you’ll often hear a proud announcement: “We hit 96 percent PM compliance this quarter.”
Applause. But here’s the question nobody asks: Did those PMs actually prevent anything?
Did the number of unplanned breakdowns drop? Did the mean time between failures improve? Did the top-ten bad actors get better? If the answer is no, then the PM program is generating activity, not reliability.
This is the “pencil-whip” problem in its most systemic form. Checklists get ticked under time pressure. Inspections are performed visually in thirty seconds instead of thoroughly in twenty minutes. The KPI measures whether the task was completed, not whether it was completed well or whether it mattered.
5. Parts and Workmanship Quality Gaps
A preventive maintenance program is only as good as the parts and the hands executing it.
In cost-conscious environments, maintenance teams are pressured to use lower-cost, non-OEM spare parts that may not meet original specifications for tolerance, material grade, or thermal rating. A seal that lasts eighteen months at OEM spec might last four months as a budget alternative.
Add to this the reality of rushed repairs during breakdowns. When a line is down and losing thousands per hour, the pressure to “just get it running” overrides the discipline of torquing to spec, following proper alignment procedures, or allowing adequate cure time. The asset goes back online, but it’s been set up for the next failure.
High technician turnover and skill gaps compound the problem further.
6. Siloed Information Between Teams
The operations team adjusts a setpoint to push throughput higher. Engineering modifies a piping run during a shutdown. Procurement switches to a different grade of hydraulic fluid.
And maintenance is never told.
The PM checklist still reflects the original operating parameters. The inspection criteria no longer match reality. The asset is being asked to perform in conditions its maintenance plan doesn’t account for. Without a single, shared source of truth for asset configuration and history, these cross-functional blind spots become breeding grounds for repeat failures.
7. No Continuous Improvement Loop
Perhaps the most fundamental gap: PM programs are treated as “set and forget.”
They are designed during commissioning or the first year of operation and then run on autopilot. There is no quarterly or semi-annual review asking:
- Is this task still relevant?
- Has the failure mode changed?
- Should this interval be shorter, longer, or replaced by a condition-based check?
- Did this asset fail again despite the PM, and if so, why?
Without tracking Mean Time Between Failures (MTBF), Mean Time To Repair (MTTR), and repeat-failure rates — and feeding those metrics back into PM design — the program calcifies. It becomes a schedule to be survived, not a strategy to be sharpened.
The Hidden Cost of Ignoring Repeat Failures
The direct costs are obvious and painful: repeated spare parts, emergency labor, overtime premiums, expedited shipping, contractor call-outs.
But the indirect costs dwarf them.
- Unplanned downtime on a critical production line can cost anywhere from a few thousand to several hundred thousand dollars per hour depending on the industry.
- Cascading failures: A failed bearing seizes a shaft, which damages a coupling, which misaligns the driven equipment. A $200 part becomes a $40,000 repair.
- Safety and compliance risk: Repeat failures on safety-critical equipment — pressure vessels, emergency shutdown systems, fire pumps — carry regulatory and human consequences that no budget line can absorb.
- Cultural erosion: When technicians spend their shifts fighting the same fires repeatedly, morale collapses. Maintenance becomes “the fix-it department” instead of a strategic function. Skilled people leave. Tribal knowledge walks out the door.
A single repeat failure on a critical conveyor in a packaging plant, for example, can cost $15,000 to $50,000 per hour in lost throughput, scrap, and downstream scheduling disruption. Multiply that by four occurrences a year, and the “small” bearing replacement becomes a six-figure problem.
Shifting from “Doing PM” to “Preventing Failures” – What Actually Works
Breaking the repeat-failure cycle doesn’t mean abandoning preventive maintenance. It means building a reliability-centered framework around it.
Layer in condition-based and predictive maintenance. Vibration analysis, thermography, ultrasonic inspection, oil analysis, and motor current signature analysis detect degradation at the P-point — often weeks or months before functional failure. These technologies complement time-based PM by telling you when to intervene, not just that the calendar says you should.
Mandate structured RCA on every repeat failure. If an asset has failed two or more times within a rolling twelve-month window, a formal root cause analysis should be non-negotiable. The output must include corrective actions that feed back into updated PM tasks, operating procedures, or design modifications.
Enforce failure coding and data discipline. Every corrective work order should capture a standardized failure code, the affected component, the probable cause, and the corrective action taken. This is the raw material for pattern recognition. Without it, you are managing maintenance by anecdote.
Optimize PM dynamically. Use actual failure data, condition monitoring trends, and operating history to adjust task frequency and content. Remove PM tasks that add no value. Add tasks that address identified failure modes. Treat the PM schedule as a living document, not a laminated poster.
Close the loop with regular reviews. Conduct quarterly or semi-annual PM effectiveness audits. Compare failure trends against PM execution. Ask: Did this PM prevent the failure it was designed to prevent? If not, redesign it.
Invest in your technicians. Training, adequate time allocation, proper tooling, and accountability for work quality — not just work speed. A technician who has forty-five minutes to do a thorough inspection will catch what a technician with fifteen minutes will miss.
How a Modern Maintenance Platform Closes These Gaps
None of the above is possible if asset history lives in filing cabinets, spreadsheets, or the memories of three senior technicians. This is where a purpose-built maintenance and asset management platform becomes the connective tissue.
A platform like TeroTAM is designed to address exactly the structural gaps described above:
- Centralized, structured asset history. Every work order, every failure, every inspection is logged against the asset with standardized failure codes. Patterns surface automatically. The fourth bearing failure is no longer a surprise; it’s a data point in a trend line.
- Condition-triggered work orders. Instead of relying solely on calendar-based scheduling, PM and inspection tasks can be triggered by condition data — vibration thresholds, temperature readings, run-hour counters — so intervention happens when the asset actually needs it.
- Built-in RCA workflows. When a repeat failure is flagged, the platform can trigger a structured root cause analysis process, ensuring findings are documented, assigned, and fed back into updated PM checklists.
- Dashboards that measure outcomes, not just activity. Track repeat-failure rates, MTBF trends, PM effectiveness ratios, and top-bad-actor rankings alongside PM compliance. Shift the conversation from “Did we do the PM?” to “Did the PM work?”
- Mobile execution with rich data capture. Technicians record inspections with photos, voice notes, and structured checklists on mobile devices, eliminating the pencil-whip problem and creating an auditable record of what was actually done.
- Cross-functional visibility. Operations, engineering, and maintenance work from the same asset record. Configuration changes, operating parameter shifts, and modification histories are visible to everyone who touches the asset.
The goal isn’t to replace the maintenance team’s expertise. It’s to give that expertise a memory, a structure, and a feedback loop so that every failure teaches the organization something permanent.
A Practical 90-Day Action Plan for Maintenance Leaders
If this article resonated and you want to start breaking the repeat-failure cycle, here is a focused, achievable plan for the next quarter.
Days 1–15: Audit your top-ten repeat offenders.
Pull work order history for the ten assets with the most corrective interventions in the last twelve months. Identify the common failure modes, components, and contributing factors.
Days 16–30: Enforce failure coding.
Implement standardized failure and cause codes on every corrective work order. Train technicians and supervisors on consistent usage. Make it a non-negotiable field before a work order can be closed.
Days 31–50: Trigger RCA on repeat failures.
Establish a policy: any asset failing two or more times in twelve months automatically triggers a structured root cause analysis. Assign ownership. Set a deadline for corrective actions.
Days 51–70: Review and prune PM tasks.
For your top-ten offenders and any other critical assets, review existing PM checklists. Remove tasks that add no measurable value. Add condition-based inspections where appropriate. Adjust intervals based on actual failure data.
Days 71–90: Shift your KPIs.
Add Repeat Failure Rate and MTBF Trend to your maintenance scorecard alongside PM compliance. Present them in your next operations review. Make the organization accountable for outcomes, not just activity.
Ninety days. Ten assets. A measurable shift from firefighting to prevention.
Summing it up
Preventive maintenance is a foundation. It is not the finished building.
For decades, the industry has equated “having a PM program” with “being proactive about reliability.” And for decades, the same pumps, motors, conveyors, and compressors have kept failing in the same ways, on the same schedules, despite the PM.
The organizations that finally break this cycle are not the ones doing more PM. They are the ones treating maintenance as a learning system. Every failure is interrogated. Every finding is codified. Every PM task is questioned, refined, and held accountable to an outcome.
The goal is not to keep the schedule full. The goal is to make the next failure impossible.
And that starts with the willingness to look at your PM program honestly, ask whether it’s actually preventing anything, and build the data, the discipline, and the tools to close the loop.
Explore how TeroTAM helps maintenance teams move from reactive checklists to reliability-driven operations.