Skip to main content
Risk ManagementFacility ManagementAsset Management

Facility Risk Assessment: Plan for Failure Before It Happens

Most facility teams know where their biggest risks are, but few have written down what failure would mean or what they would do about it. Here is a simple way to do it, system by system.

October 9, 2026
•
9 min read
•Asset Management

One of the most overlooked tasks in facilities management is working out where your biggest risks are.

For much of my career in maintenance, I kept that information in my head. I knew which systems worried me, but I never wrote down what could go wrong or what we would do about it. The system I was using did not even have a place to capture it. If I had been away the night the boiler failed, whoever got the call would have been starting from nothing.

In my opinion, this should be one of the first things you do as a new facility manager. This post sets out how: which systems to start with, the four questions to ask of each one, a worked example, and a one-page template you can copy.

The short answer: For each critical system, walk through a failure scenario and write down four things: what happens if it goes down, how long you can manage without it, what you would do (and who you would call), and what you can do now to reduce the risk.

Here is what this post covers:

  • What a facility risk assessment is, and the half most teams skip
  • Which building systems to assess first
  • The four questions that turn a worry into a response plan
  • A worked example and a copy-ready failure scenario template
  • How often to review the plan, and what should trigger a review

What is a facility risk assessment?

A facility risk assessment identifies which building systems would cause the most harm if they failed, how likely that failure is, and what you would do about it. It has two halves, and most teams only do the first.

  • The risk score ranks your systems. The usual method multiplies consequence of failure by likelihood of failure, each on a 1 to 5 scale, for a score from 1 to 25. It tells you where to spend maintenance time and capital first.
  • The failure scenario is the plan for when a system actually fails: the impact, the tolerable downtime, who to call, what you need, and what you can do now to make it less likely or less painful.

The score is for budget meetings. The scenario is for 2 AM. Risk-based standards such as ISO 55000 expect both: you identify the risks to your assets and you plan how to treat them.


Why does risk knowledge stay in one person's head?

Because there is rarely anywhere else to put it. A typical CMMS holds assets, work orders and PM schedules. It has no field for "if this fails in January we have six hours" or "the contractor who knows this unit retired, call this one instead." So that knowledge lives with the person who has been there longest.

That works until it does not:

  • People leave. When a long-serving operator retires, the risk knowledge retires with them.
  • Failures happen off-shift. The person on call is often not the person who knows the system best.
  • Nothing gets funded. A worry in your head cannot go into a capital request. A documented scenario with a risk score can.
  • Nobody prepares. The spare part, the rental agreement and the vendor number are only arranged in advance if someone wrote down that they would be needed.

Which building systems should you assess first?

Start with the systems whose failure would hurt someone, close the building or break a regulation. For most facilities that is a list of five to ten:

  • Heating plant: boilers, pumps, and the controls that run them
  • Cooling and refrigeration: chillers, ice plants, cold storage
  • Electrical service: main switchgear, transformers, distribution
  • Emergency power: generator, transfer switch, UPS
  • Fire alarm and sprinkler systems
  • Domestic water and backflow prevention
  • Elevators, especially where they are the only accessible route
  • Roof and building envelope

Then add anything with a single point of failure (one boiler, one pump, one feed) or a long lead time for replacement parts. A system you cannot repair for twelve weeks deserves a plan even if it rarely fails.


How do you run a failure scenario? The four questions

Take one system at a time. Ideally, walk to it with the person who knows it best, and answer these four questions out loud before writing anything down.

1. What happens if this system goes down?

Describe the consequences in plain terms, then check them against six kinds of impact: safety, service disruption, financial loss, regulatory, environmental and reputation. A chiller failure at an office is an uncomfortable afternoon. The same failure at an ice rink during a tournament is cancelled bookings and a refund problem. Be specific about who is affected.

2. How long can we manage without it before it becomes serious?

This is your tolerable downtime, and it is the number most teams have never put a figure on. It almost always depends on the season: a boiler outage in July can wait for parts; in January it may give you hours. Write down both. This single number tells you whether your response is "schedule a repair" or "everyone drops what they are doing."

3. What would we do, who would we call, and what would we need?

This is the response plan. Write the first three actions in order. List the vendor and their after-hours number, your contract or account number, and who in your organization can approve emergency spending or decide to close the building. Then list what the response depends on: spare parts, rental equipment, isolation valve and shutoff locations, drawings, keys and access.

4. What can we do now to reduce that risk?

This is where the exercise pays for itself. Answers usually fall into two groups. Some reduce the likelihood of failure: better preventive maintenance, a condition assessment, or a planned replacement. Others reduce the impact: stocking the parts that fail most, arranging a standby rental, installing a connection point for a temporary unit, or training a second person. Each one becomes a work order, a purchase or a line in the capital plan.

Rule of thumb: If the answer to question 2 is shorter than the time it takes to get the answer to question 3, the gap is your real risk, and question 4 is how you close it.


Worked example: the boiler plant in winter

Here is what a finished scenario looks like for a heating plant in a Canadian community facility. Note the risk score at the top: the plan sits next to it, so anyone looking at the most critical system also sees what to do about it.

Failure scenario
Heating plant: Boilers 1 and 2
System: Heating · Main building
20
Critical · CoF 5 × LoF 4
If it goes down
No heat in the building. Pipes at risk in exterior walls and the sprinkler riser room. Programs cancelled; members sent home.
How long we can last
About 6 hours at -15 °C before indoor temperatures fall below safe levels. 2 to 3 days in shoulder season.
Who we call
Mechanical contractor after-hours line (contract on file). Gas utility for supply issues. Insurer if there is water damage.
What we need
Igniter and flame sensor on the shelf. Temporary boiler rental with a pre-agreed hookup point. Isolation valve locations on the floorplan.
What we do now
Stock the two common failure parts. Sign a standby rental agreement. Move Boiler 2 inspection to September. Add the plant to next year’s capital request.
Illustrative. Your tolerable downtime depends on your climate, your building envelope and who uses the space.

Look at the last row. Four of the five actions cost little or nothing, and none of them would have happened if the scenario had stayed in someone's head. The fifth, the capital request, is now backed by a documented consequence and a tolerable downtime of six hours, which is a much easier case to make to a board than "the boiler is old."


A failure scenario template you can copy

Keep each scenario to one page. If it does not fit, it will not be read during an outage.

Failure scenario template
FieldWhat to write
System and assetsWhich system, and which assets are part of it?
ConsequenceSafety, service, financial, regulatory, environmental, reputation
Tolerable downtimeHours or days before it becomes serious, by season
First responseThe first three things to do, in order
Who to callVendor, after-hours number, contract or account number
What you needSpares, rentals, shutoff locations, drawings, access
Who decidesWho can close the building or approve emergency spend
Mitigation nowWhat reduces the likelihood or the impact this year
Last reviewedDate and name

Fill it in for your top five systems first. A rough scenario written this month is worth far more than a perfect one planned for next year.


Risk in your head vs risk you have documented

Attribute
In someone's head
Documented
When the expert is awayWhoever is on call starts from scratchThe plan is on the asset, findable by anyone
Tolerable downtimeA feeling that it is "pretty bad"A number, by season
Vendor and partsSearched for during the outageListed, and spares stocked in advance
Capital requests"It is old and I am worried about it"Risk score plus a written consequence
Staff turnoverKnowledge leaves with the personKnowledge stays with the facility
MitigationGood intentionsWork orders, purchases and capital items

The same knowledge, before and after it is written down against the asset.


How often should you review failure scenarios?

At least once a year, and a review before each system's critical season is better: heating in the fall, cooling in the spring. Also review a scenario whenever one of these happens:

  • The system actually fails. Compare what happened with what you wrote.
  • The system is replaced, modified or expanded.
  • A vendor, contract or key staff member changes.
  • The use of the space changes, such as a new tenant or program.

Phone numbers and part numbers go stale faster than anything else in the plan, so check those every time.


Where AssetLab fits

When I built AssetLab, I made sure there was a place to document this information next to the asset it belongs to, and to make it easy to find and act on. The risk management module scores every asset by consequence of failure times likelihood of failure. Consequence is assessed across the same six impacts used above, and likelihood is drawn from the age, condition and maintenance history your team already records.

Alongside that score, you record the failure scenario and response plan. So the asset at the top of your risk ranking is also the one with the plan attached, and the mitigation work goes straight into work orders, PM schedules and the replacement planner. Risk scoring is part of the AssetLab 360 and Enterprise plans.

How AssetLab helps

How AssetLab handles risk

Risk in AssetLab is not a separate spreadsheet. It reads the assets, inspections and work orders you already keep, and keeps the plan where the asset is.

5×5 risk matrix

Every asset scored 1 to 25 by consequence × likelihood of failure.

Six consequence dimensions

Safety, service, financial, regulatory, environmental and reputation.

Data-driven likelihood

Drawn from age, condition scores and maintenance history.

Scenarios and response plans

Document what failure means and what to do, next to the risk score.

Reusable risk profiles

Set consequence once per system and every asset in it inherits it.

From risk to action

High-risk assets feed work orders, PM schedules and the capital plan.


Frequently Asked Questions

What is a facility risk assessment?

A facility risk assessment identifies which building systems would cause the most harm if they failed, how likely that failure is, and what the team would do about it. A complete one has two halves: a risk score (consequence of failure multiplied by likelihood of failure) that ranks the systems, and a written failure scenario for each critical system that records the impact, how long you can operate without it, who to call, what you would need, and what you can do now to reduce the risk.

What questions should a failure scenario answer?

Four. What happens if this system goes down? How long can we manage without it before the consequences become serious? What would we do, who would we call, and what would we need? What can we do now to reduce that risk? The answers become the response plan, and the last one becomes a list of work orders, purchases and capital requests.

Which building systems should a new facility manager assess first?

Start with the systems whose failure would hurt people, close the building or break a regulation: heating plant and boilers, cooling and refrigeration, electrical service and switchgear, emergency power, fire alarm and sprinklers, domestic water, elevators, and the roof. Five to ten systems is enough for a first pass. Add anything with long-lead replacement parts or a single point of failure.

What is the difference between a risk score and a response plan?

A risk score tells you which systems matter most and in what order. A response plan tells you what to do when one of them fails. The score is for prioritizing maintenance and capital; the plan is for the night the boiler trips. You need both, because a ranked list with no plan still leaves the team improvising, and a plan with no ranking tends to get written for the wrong systems.

How often should failure scenarios be reviewed?

At least once a year, and also after any real failure, after a system is replaced or modified, when a key vendor or staff member changes, and before the season when the system is most critical, such as heating before winter. Phone numbers and part numbers go stale faster than anything else in the plan.

Can a CMMS store failure scenarios and response plans?

Many cannot: they hold assets, work orders and PM schedules, but have nowhere to record what failure would mean or what to do about it. AssetLab scores every asset by consequence and likelihood of failure and gives you a place to document the failure scenario and response plan alongside that score, so the plan is found where the asset is, not in a binder. Risk scoring is part of the AssetLab 360 and Enterprise plans.


Write it down before you need it

Every facility manager carries a list of systems that worry them. The difference between a worry and a plan is an hour per system and four questions: what happens, how long you can last, what you would do, and what you can do now. Have you documented your risks and response plans, or is most of that knowledge still in someone's head?

Pick your top five systems. Answer the four questions. Put the answers where the next person will find them.

If you want to see risk scoring and response plans on your own assets, we can walk you through it in about 20 minutes.

A maintenance technician with a tablet inspecting pumps in an industrial plant

Built for the people who make things happen.