Select Page

Defect elimination is the structured process of permanently removing the root causes of recurring failures rather than repeatedly repairing their symptoms. It runs as a standing site programme that ranks failures by business impact, investigates the highest-cost ones through root cause analysis, implements the fix, and then verifies against the site’s own failure data that the failure mode actually stopped.

HolisticAM runs defect elimination as a monthly cycle for operating sites. This guide covers the loop, the data it depends on, how it relates to root cause analysis, and what the arithmetic looks like on a recurring failure.

What is defect elimination?

A defect, in this context, is any condition that causes an asset to fail to do what it is meant to do. That covers the obvious physical ones, a bearing that runs hot or a seal that leaks, and the ones that never appear in a maintenance system: a task written against the wrong failure mode, a spare specified incorrectly, an operating procedure that loads a machine outside its design envelope.

Defect elimination is the programme that finds those conditions, ranks them by what they cost the business over a year, removes the cause, and proves the removal worked. The word doing the work in that sentence is permanently. Repairing a failure faster is maintenance. Making the failure stop happening is defect elimination.

The practical test of whether a site has a real programme is simple. Almost every operation has a defect elimination process on paper. Ask instead whether the failure rate has moved. If nobody can answer that from data, what exists is an investigation process, not a defect elimination programme.

The defect elimination loop

The programme runs as a closed loop on a fixed cadence, typically monthly, whether or not anything significant has gone wrong recently. Five steps.

  1. Identify and log. Pull recurring failures from work order history, downtime accounting and operator reports. The unit of interest is the failure mode on an asset, not the individual work order.
  2. Rank by annualised business impact. Cost each failure mode across a full year, including production loss, secondary damage and labour, not just the repair line. This step is where most programmes go wrong and it is covered below.
  3. Analyse the cause. Take the top-ranked items into a structured root cause analysis. Investigate to the point where the causes are evidenced, not to the point where someone is satisfied with an answer.
  4. Implement and verify. Turn causes into owned, dated actions. Then go back to the failure data at three and six months and check whether the failure rate actually changed. An action marked complete is not a verified fix.
  5. Lock it in. Fold the fix into the maintenance strategy, the design standard, the spare specification or the operating procedure, so it survives the people who made it. A fix that lives only in a closed action will be undone by the next component swap.

Step four is the one that separates programmes that change the failure rate from programmes that generate paperwork. Skipping it means the same failure reappears six months later and is treated as a new event, because nobody went back to the data.

Where the data comes from

Defect elimination is a data-fed process and it draws from three sources, each with a characteristic weakness.

  • Work order history. The primary source and the one that tells you frequency. Its weakness is coding: failure modes recorded inconsistently, or not at all, so the same recurring problem appears as a dozen unrelated jobs.
  • Downtime accounting. Tells you consequence, which is what makes ranking possible. Its weakness is attribution, because downtime is often booked to the area rather than the asset, and almost never to the failure mode.
  • Operator reports and shift logs. The richest source for the failures that never generate a work order at all, and the one most often ignored. A stoppage cleared by an operator in fifteen minutes leaves no maintenance record and can still be the largest annual loss on the site.

Why poor work order data cripples the programme

Ranking is the whole value of defect elimination. It is what stops a site from spending its improvement effort on the failure that made the most noise instead of the failure that cost the most money. Ranking depends entirely on being able to group work orders by failure mode and attach a consequence to each group.

When failure coding is inconsistent, that grouping cannot happen. Fourteen occurrences of one problem look like fourteen unrelated small jobs, none of which is large enough to attract attention. Meanwhile the single dramatic event that shut the plant for a day gets the investigation, gets the recommendations, and consumes the improvement capacity for the quarter.

A twenty minute stoppage that happens fourteen times a year usually costs more than the single large event that got the investigation. Nobody has ever added it up, because the work orders were never grouped.

This is why a defect elimination programme normally starts with an honest assessment of whether the maintenance history can support ranking at all. Where it cannot, the first phase is data remediation, and that should be scoped and said out loud rather than discovered three months in. A programme built on unrankable data will produce a register, activity and reports, and will not change the failure rate.

ISO 14224 is the reference worth knowing here. It sets out how reliability and maintenance data should be collected and classified for equipment in the process industries, including failure mode taxonomies. Sites that adopted its structure for failure coding can rank; sites that let coding evolve informally usually cannot.

How defect elimination uses root cause analysis

Root cause analysis is the investigative step inside the loop. Its job is to establish, with evidence, the chain of causes that produced a failure, so that the action list addresses causes rather than symptoms.

HolisticAM facilitates investigations using the Apollo Root Cause Analysis methodology, a cause-and-effect approach that maps every causal branch and requires each cause-effect link to be verified rather than asserted. Our instructors are accredited Apollo RCA facilitators, accredited by Apollonian Publications.

The reason the method choice matters for defect elimination specifically is that a linear technique stops at a single causal chain. Recurring failures rarely have one cause. A seal fails because of shaft deflection, and because the flush plan was wrong for the duty, and because the PM task inspected the wrong thing. Address one and the failure returns. A cause-and-effect map surfaces all three, and the action list can then be prioritised across them.

A worked example: the pump seal that fails fourteen times a year

The following is an illustrative composite built from patterns common across mineral processing operations. It is not drawn from any single client engagement.

A slurry pump gland seal on a processing circuit fails fourteen times a year. It is not a dramatic failure. Each event is cleared in a shift and nobody escalates it.

The repair-cost view. This is what the maintenance system reports, because it is what the work orders capture.

  • Labour: two fitters, six hours each event, at $95 per hour = $1,140
  • Seal kit and consumables = $850
  • Per event: $1,990. Annualised across fourteen events: about $27,900

Twenty-eight thousand dollars a year is a rounding error on a processing budget. Ranked on this number, the failure never reaches an improvement register.

The total business impact view. This is what the failure actually costs.

  • Production loss. Each failure takes the circuit down for three and a half hours once isolation and restart are counted. Surge capacity absorbs part of it, leaving roughly ninety minutes of genuinely lost throughput at 400 tonnes per hour and a $18 per tonne margin, so about $10,800 per event, or $151,200 a year
  • Secondary damage. Roughly every third failure destroys the shaft sleeve, at $4,200 a time, adding about $19,600 a year
  • Expedited freight on the four occasions a year the spare was not held, at $900 each: $3,600
  • Plus the repair cost above: $27,900
$202,300
Illustrative annual cost of one recurring seal failure once production loss and secondary damage are counted, against the $27,900 the work orders report. A factor of seven.

What a permanent fix looks like. The investigation finds three verified causes: a bearing housing worn beyond tolerance allowing shaft deflection, a seal flush plan carried over from a different duty, and a PM task inspecting seal condition rather than the shaft runout that was driving the failure. The remedy is to re-machine and re-bore the housing, change the flush plan to suit the actual slurry duty, and rewrite the PM task to measure runout at an interval derived from the observed degradation rate. Roughly $18,000 once, against $202,300 a year.

The point of the example is not the arithmetic. It is that the failure was invisible to the ranking process for as long as ranking was done on repair cost, which is the default in most maintenance systems.

Find out what your recurring failures actually cost

A Reliability Assessment puts a dollar figure on the failure modes costing your site the most, ranked on annualised business impact rather than repair cost. Fixed scope, fixed fee. Most sites are surprised by what sits at the top of the list.

Book a Reliability Assessment Talk to a reliability engineer

Defect elimination vs root cause analysis: what is the difference?

These two get used interchangeably and they are not the same thing. Root cause analysis is a method. Defect elimination is a programme that uses it.

Root cause analysis Defect elimination
What it is An investigation method A standing programme
Trigger A specific event, usually a significant one A fixed monthly cadence, event or not
Scope One failure The whole ranked population of recurring failures
Selection Chosen by consequence or by who escalated it Chosen by annualised cost, which usually surfaces different failures
Output A verified cause chain and a recommendation list Closed actions, a revised strategy, and measured change in failure rate
Ends when The investigation is complete It does not end. It runs as a loop
Success measure Causes evidenced rather than asserted The failure mode stopped, verified in the site’s own data

RCA without defect elimination produces a drawer of good investigations whose recommendations were never closed. Defect elimination without competent RCA produces a register of actions that address symptoms, closed on schedule, with the failure rate unchanged. Sites often have one and assume they have both.

Frequently asked questions

How do you prioritise which defects to eliminate first?

Rank by annualised business impact, meaning frequency multiplied by the full consequence of each occurrence: production loss, secondary damage, expedited spares and labour. Ranking on repair cost alone systematically hides high-frequency, low-repair-cost failures, which are usually where the largest recoverable losses sit.

How many defects should a site work on at once?

Fewer than most sites attempt. Capacity to close actions, not capacity to investigate, is the binding constraint. A register that grows faster than it closes is the signal to reduce the number in progress, because once people stop expecting actions to close, attendance and data quality both fall away.

What roles does a defect elimination programme need?

A facilitator competent in the analysis method, a person accountable for the register who is senior enough to chase closure across departments, and named owners with dates for each action. Reliability engineering provides the analysis and the verification. The critical role is register ownership, and it is the one most often left unassigned.

How long before a defect elimination programme pays back?

The first ranked list often identifies more annualised loss than the programme costs, but that is potential rather than realised. Realised savings depend on closure and verification, so expect the first measurable change in failure rate at three to six months on the earliest items, provided actions are actually closing.

What data do you need to start?

Work order history with failure coding consistent enough to group occurrences by failure mode, downtime records attributable to assets, and access to operator logs for the failures that never raise a work order. Where failure coding will not support grouping, data remediation is the first phase and should be scoped explicitly rather than discovered later.

Is defect elimination the same as continuous improvement?

No. Continuous improvement is a broad philosophy applied to any process. Defect elimination is narrower and more testable: it targets the causes of asset failures specifically, ranks them on financial consequence, and measures success as a verified change in failure rate rather than as improvement activity completed.