“We need a reliability program” is usually said in the week after a bad failure, and the next move is almost always wrong: hire a reliability engineer, buy some software, announce a program. Eighteen months later the engineer has become a firefighter, the software holds a criticality study nobody uses, and the program is a slide from last year’s budget cycle. The sequence below is the one that survives contact with an operating mine site. It is built from the engagements where we have stood the function up and, just as instructively, the ones where we were called in after a first attempt stalled.
HolisticAM (Holistic Asset Management) is an Australian reliability engineering and asset management consultancy that builds and runs reliability capability for mining and heavy industry, from facilitated studies to a fully managed function. This is the playbook we actually use.
Key takeaways:
- Start with evidence, not appointments: your own failure and cost data tells you where the program should aim before anyone is hired or trained.
- Build the smallest working loop first (track failures, analyse the worst, fix one, prove it stayed fixed), then scale what works.
- Foundations before headcount: in our experience, a lone reliability engineer hired into a site with no working system becomes a firefighter within months.
In This Guide
- Before Anything Else: Know What the Program Is For
- Step 1: Read Your Own Evidence First
- Step 2: Rank the Assets
- Step 3: Build the Smallest Working Loop
- Step 4: Connect the Loop to the Weekly Machine
- Step 5: Now Decide the Capability Model
- The First Ninety Days, Roughly
- Why Reliability Programs Fail
- Frequently Asked Questions
- Further Reading
Before anything else: know what the program is for
If the program needs a board-level case before it gets funded, start with The Reliability P&L Manifesto.
A reliability program is not a department, a software purchase or a set of ceremonies. It is a working loop: failures get recorded, the worst get analysed, the causes get eliminated or managed by a better strategy, and the results get measured. Everything else (roles, tools, training) exists to run that loop at scale. Start by writing one sentence naming the loss the program exists to reduce: unplanned downtime on the crushing circuit, repeat failures on the mobile fleet, a ramp-up that keeps slipping. A program aimed at everything improves nothing.
Step 1: read your own evidence first
Before appointing anyone, pull what the site already knows: work order history, downtime records, the biggest failures of the last year, and where the maintenance money went. The data will be imperfect; it is still the map. You are looking for three things: where the losses concentrate, which failures repeat, and whether the records can support analysis at all. This is a bounded piece of work with a defined output, and it is exactly what a fixed-scope Reliability Assessment does: reads the data, names the failure modes driving the losses, and hands you a costed next step you own.
The evidence step also settles the argument the program will otherwise have every month: what to work on first. Loss data ends that debate before it starts.
Step 2: rank the assets so effort lands where it matters
An asset criticality assessment (ranking equipment by the consequence of its failure) is the cheapest piece of reliability engineering a site can do, and it steers everything after it: which assets get deep strategy work, which failures earn investigation, which spares matter. Without it, programs default to working on whatever failed most recently, which is how a year of effort gets spent on loud-but-cheap problems while the quiet expensive ones continue.
Step 3: build the smallest working loop
Resist the program-launch instinct. Instead, get one full cycle working on a handful of assets:
- Track failures properly on the chosen assets. Coded against a usable standard such as ISO 14224, dated, with downtime attributed. If close-out quality is poor, fix it for these assets first; the loop dies without its data.
- Analyse the worst recurring failure. Structured root cause analysis on one failure that hurts (our guide to root cause analysis covers method selection), or Weibull analysis where the history supports it.
- Fix it, and check the strategy. Eliminate the cause where possible; where not, make sure the maintenance strategy addresses the failure mode (RCM logic per SAE JA1011, applied at whatever depth the asset earns).
- Prove it stayed fixed. Repeat-failure count on that asset, before and after. One demonstrated win buys the program more support than any launch deck.
This is defect elimination in miniature, and it is deliberately small: the point is to prove the site can run the loop before scaling it.
Step 4: connect the loop to the weekly machine
Reliability wins are delivered through the maintenance system, not around it. The strategy changes the loop produces must land in the CMMS as tasks, the tasks must be planned and scheduled properly, and close-out data must flow back. If the site’s planning and scheduling is broken, the reliability program’s outputs evaporate at execution, so the program and the work-management basics improve together or not at all. A short set of maintenance and reliability KPIs on consistent definitions, baselined before targets, keeps the whole thing honest.
Step 5: now decide the capability model
Only after the loop is proven does the headcount question have a good answer, and the counterintuitive part is the order. A reliability engineer hired first, into a site with no working loop, no criticality basis and no protected role, becomes an expensive firefighter; we have watched it happen often enough to call it the default outcome. Hire into a working system instead, or run the function as a managed capability while the system matures. The full decision (in-house hire, consultant campaigns, or a managed reliability service, and how they combine) is covered in in-house vs consultant vs managed. For sites that want the function running without waiting on recruitment, that is exactly what our Reliability Operating Centre provides: remote analysis with onsite days built in.
The first ninety days, roughly
- Weeks 1 to 4: evidence read, one-sentence aim written, criticality ranking drafted, target assets chosen.
- Weeks 5 to 8: failure tracking fixed on target assets; first structured analysis run; first fixes raised through the CMMS.
- Weeks 9 to 13: first repeat-failure evidence in; loop reviewed; scale decision made (more assets, deeper strategy work, capability model settled).
Sites differ and the calendar flexes; the sequence does not. Evidence, ranking, loop, connection, then capability.
Why reliability programs fail
The post-mortems repeat themselves:
- Started with headcount or software instead of evidence. The program had a face and a licence before it had an aim.
- Aimed at everything. No criticality basis, so effort followed noise.
- The loop never closed. Analyses were done and recommendations written; nothing changed in the CMMS, and nobody checked whether failures recurred.
- The engineer became a firefighter. Urgent ate important within a quarter.
- Measured activity instead of results. Studies completed and meetings held, while repeat failures and unplanned downtime never moved.
Every one of these is avoidable, and every one is cheaper to avoid than to unwind.
Starting with HolisticAM
HolisticAM builds reliability capability for mine sites in whatever form the site needs: a fixed-scope Reliability Assessment to read your evidence and name the starting point, facilitated criticality and strategy work, structured defect elimination, or the full function run through the Reliability Operating Centre while your system and team mature. Evidence first, fixed scope first, headcount when the system deserves it.
Frequently asked questions
What is the first step in starting a reliability program?
Read your own evidence before appointing anyone: work order history, downtime records and the year’s worst failures. The data shows where losses concentrate, which failures repeat, and whether your records can support analysis. Pair that with an asset criticality ranking and you have the program’s aim and its first targets on evidence rather than opinion. Hiring, software and training all come after, shaped by what the evidence says.
What does a mine site need before starting a reliability program?
Three foundations: a CMMS with work order history the program can read (imperfect is normal; absent means fixing data capture first), a criticality ranking so effort lands on assets whose failure hurts most, and leadership willing to protect the program’s time from daily firefighting. Money and headcount matter less at the start than most sites assume; a small proven loop beats a large announced program.
Why do most reliability programs fail?
The recurring causes: starting with headcount or software instead of evidence, aiming at everything because no criticality basis existed, analyses that never changed anything in the CMMS, the reliability engineer dissolving into firefighting for lack of a protected role, and measuring activity (studies, meetings) instead of results (repeat failures, unplanned downtime). All five are sequencing failures, which is why the starting order matters more than the budget.
How long does it take for a reliability program to show results?
The first credible evidence, a repeat failure eliminated and shown to stay eliminated on a target asset, is achievable inside the first ninety days if the program starts small. Site-level movement in unplanned downtime and planned-work ratio builds over quarters as strategy changes work through the plan. Programs that promise plant-wide transformation in a quarter are usually measuring activity; programs that show one asset’s failures actually stopping are building something real.
Who should own the reliability program on site?
Ownership needs two layers. A senior operational owner (typically the maintenance or asset manager) who protects the program’s time, chairs its rhythm and owns its results, and a technical lead who runs the analysis and the loop, whether that is an in-house engineer, a consultant presence or a managed function. What fails is ownership by committee, or a technical lead with no senior cover; the program’s time gets eaten within a quarter.