Maintenance KPIs exist to change decisions, and that single test sorts the useful ones from the wallpaper. MTBF, MTTR, availability and OEE are the four metrics every maintenance dashboard carries, yet on many sites nobody can say what decision any of them changed last month. This guide explains what each metric actually measures, how to calculate it, what it is genuinely good for, and the traps that turn measurement into theatre.
HolisticAM (Holistic Asset Management) is an Australian reliability engineering and asset management consultancy that builds maintenance KPI frameworks and reporting for heavy industry, on standard definitions, wired to the decisions each number is supposed to drive. Everything below reflects how we set these metrics up on operating sites.
Key takeaways: MTBF measures how often equipment fails; MTTR measures how long repairs take; availability combines the two into the fraction of time the asset can run; OEE multiplies availability by performance and quality to expose total production losses. All four are lagging indicators. They tell you the score, not how to change it, which is why they need leading work-management measures alongside them and failure analysis beneath them.
In This Guide
- The Test Every KPI Must Pass
- MTBF: Mean Time Between Failures
- MTTR: Mean Time to Repair
- Availability: The Two Combined
- OEE: Overall Equipment Effectiveness
- Lagging Scores Need Leading Measures
- Why This Article Publishes No Benchmark Table
- Building a KPI Framework That Earns Its Keep
- Frequently Asked Questions
- Further Reading
The test every KPI must pass
Before the formulas: a metric earns its place on a report only if someone, looking at it, would act differently at different values. Who reads this number, and what do they do when it moves? If there is no answer, the metric is decoration. The strongest maintenance KPI sets we see are short, defined once (the SMRP metrics and the European EN 15341 standard exist precisely so sites stop inventing private definitions), baselined before targets are set, and reviewed with the reasons behind the movement, not just the movement.
MTBF: mean time between failures
What it is: the average operating time between failures of a repairable asset.
Formula: MTBF = total operating time ÷ number of failures.
Worked example (hypothetical): a slurry pump runs 4,380 hours in a half-year window and suffers 6 functional failures. MTBF = 4,380 ÷ 6 = 730 hours. The number says the pump fails, on average, about once a month of continuous running.
What it is good for: trending a specific asset or fleet against itself over time, and flagging where failures are accelerating. Rising MTBF after a strategy change is evidence the change worked.
The trap: averaging across different things. An MTBF blended across mixed failure modes, or across dissimilar assets, hides exactly the signal you need. Six failures made of three seal leaks, two bearing failures and an electrical trip are three different problems wearing one average. MTBF also says nothing about the pattern of failure: whether the asset is wearing out or failing randomly. The moment MTBF becomes an argument about replacement intervals, you have outgrown it, and the right tool is Weibull analysis of the separated failure modes.
MTTR: mean time to repair
What it is: the average time to restore the asset to service after a failure.
Formula: MTTR = total repair downtime ÷ number of failures.
Worked example (hypothetical, continuing): those 6 failures cost 48 hours of repair downtime in total. MTTR = 48 ÷ 6 = 8 hours.
What it is good for: exposing maintainability and response problems: access, spares availability, isolation complexity, diagnostic time. A rising MTTR with a stable MTBF is a supportability problem, not a reliability problem, and it points at different fixes (kitting, spares holdings, work instructions) than a falling MTBF does.
The trap: definition drift. Does the clock start at failure, at notification, or at trade arrival? Does it include waiting for parts? Any answer can work; changing the answer between months cannot. Write the definition down once and hold it.
Availability: the two combined
What it is: the fraction of required time the asset is capable of running.
Formula (one standard form): Availability = MTBF ÷ (MTBF + MTTR), or equivalently uptime ÷ (uptime + downtime).
Worked example (hypothetical, continuing): 730 ÷ (730 + 8) = 98.9%. The same result comes from the raw hours: 4,380 ÷ (4,380 + 48).
What it is good for: the bridge between maintenance performance and production planning. Availability is the number operations actually feels.
The trap: a high availability can coexist with an expensive, chaotic maintenance operation; it says nothing about what the uptime cost to achieve. It also moves slowly. A site can burn a quarter celebrating 98.9% while the failure count quietly doubles on one critical asset, so availability needs MTBF and MTTR reported beside it, not instead of it.
OEE: overall equipment effectiveness
What it is: the fraction of theoretical production capacity actually delivered, combining three losses.
Formula: OEE = Availability × Performance × Quality, where performance is actual rate versus design rate while running, and quality is the share of output that is right first time.
Worked example (hypothetical): a processing line runs at 90% availability, at 85% of design rate while running, with 98% of output in specification. OEE = 0.90 × 0.85 × 0.98 = 75.0%.
What it is good for: showing where the biggest loss actually lives. In the example, the availability number looks like the headline, but the performance loss is costing more than the downtime is. OEE’s whole value is forcing the three loss categories into one honest conversation between maintenance and production.
The trap: OEE is the most gamed metric in industry. Redefine “planned production time” generously, exclude the right stoppages, and OEE climbs while output does not. It is also meaningless as a comparison between different plants with different definitions. Use it inward, against your own baseline, with definitions nobody is allowed to quietly improve.
Lagging scores need leading measures
All four metrics above are lagging: they report the consequences of decisions made months ago. A KPI framework that can actually steer a maintenance operation pairs them with the leading work-management measures that predict them: schedule compliance and break-in rate (is the site running its plan?), planned-versus-reactive work mix, backlog health, and PM program compliance. Those measures, and the weekly routine that reviews them, belong to maintenance planning and scheduling, and they move weeks before MTBF does.
The other half of the pairing is depth beneath the KPIs. When a lagging number moves the wrong way, the follow-up questions are analytical: which failure modes drove it (failure coding and history quality), whether the pattern is wear-out or random (Weibull), and whether the maintenance strategy behind the asset still matches its failure behaviour (RCM, or a maintenance task optimisation review). A KPI without a follow-up question attached is a scoreboard, not an instrument.
Why this article publishes no benchmark table
Deliberately. Published “world-class” targets for these metrics vary wildly with definitions, industries and duty, and a target imported without a baseline mostly produces creative reporting rather than better maintenance. The pattern that works: define each metric once on a public standard (SMRP or EN 15341), baseline your own operation for a few months, then set improvement targets against your own numbers, with the definition frozen. If a consultant opens with a benchmark table instead of your data, ask what decision it is supposed to change.
Building a KPI framework that earns its keep
Keeping these measures current is one of the four services HolisticAM runs through its Reliability Operating Centre.
HolisticAM builds maintenance KPI frameworks on standard definitions, wired into reporting with agreed break-in rules and review routines, as part of our maintenance management and reliability engineering services. The starting point is the same as ever: your own data first. A Reliability Assessment reads your maintenance and failure history, names the failure modes driving your losses, and hands you a costed next step. Fixed scope, fixed fee, agreed before we start.
Frequently asked questions
How do you calculate MTBF?
Divide total operating time by the number of failures in that period: MTBF = operating hours ÷ failure count. For example, an asset that ran 4,380 hours with 6 failures has an MTBF of 730 hours. Calculate it per asset and, ideally, per failure mode; blending different assets or different failure mechanisms into one average hides the signal the metric exists to show.
What is the difference between MTBF and MTTR?
MTBF (mean time between failures) measures how often an asset fails: average operating time between failures. MTTR (mean time to repair) measures how long restoration takes: average downtime per repair. They diagnose different problems. A falling MTBF is a reliability problem in the asset or its maintenance strategy; a rising MTTR is a supportability problem in spares, access, information or response.
What is the difference between availability and reliability?
Reliability is the probability an asset performs its function without failure over a defined period; MTBF is one simple indicator of it. Availability is the fraction of required time the asset is capable of running, combining how often it fails with how long repairs take: MTBF ÷ (MTBF + MTTR). An asset that fails often but is repaired in minutes can show high availability with poor reliability.
How do you calculate OEE?
Multiply three factors: OEE = availability × performance × quality. Availability is running time against planned production time, performance is actual output rate against design rate while running, and quality is right-first-time output as a share of total output. For example, a line at 90% availability, 85% performance and 98% quality has an OEE of 75%. The value is in the three factors separately, which locate the biggest loss.
Which maintenance KPIs should a site track first?
Start with a lagging-and-leading pair on consistent definitions: availability with MTBF and MTTR beneath it for the score, and schedule compliance with break-in rate for whether the site is running its plan. Baseline them for several months before setting any targets, and attach an owner question to each: who acts on this number, and what do they do when it moves. A short set that changes decisions beats a dashboard that changes nothing.