One of the most overlooked tasks in facilities management is working out where your biggest risks are.
For much of my career in maintenance, I kept that information in my head. I knew which systems worried me, but I never wrote down what could go wrong or what we would do about it. The system I was using did not even have a place to capture it. If I had been away the night the boiler failed, whoever got the call would have been starting from nothing.
In my opinion, this should be one of the first things you do as a new facility manager. This post sets out how: which systems to start with, the four questions to ask of each one, a worked example, and a one-page template you can copy.
The short answer: For each critical system, walk through a failure scenario and write down four things: what happens if it goes down, how long you can manage without it, what you would do (and who you would call), and what you can do now to reduce the risk.
Here is what this post covers:
- What a facility risk assessment is, and the half most teams skip
- Which building systems to assess first
- The four questions that turn a worry into a response plan
- A worked example and a copy-ready failure scenario template
- How often to review the plan, and what should trigger a review
What is a facility risk assessment?
A facility risk assessment identifies which building systems would cause the most harm if they failed, how likely that failure is, and what you would do about it. It has two halves, and most teams only do the first.
- The risk score ranks your systems. The usual method multiplies consequence of failure by likelihood of failure, each on a 1 to 5 scale, for a score from 1 to 25. It tells you where to spend maintenance time and capital first.
- The failure scenario is the plan for when a system actually fails: the impact, the tolerable downtime, who to call, what you need, and what you can do now to make it less likely or less painful.
The score is for budget meetings. The scenario is for 2 AM. Risk-based standards such as ISO 55000 expect both: you identify the risks to your assets and you plan how to treat them.
Why does risk knowledge stay in one person's head?
Because there is rarely anywhere else to put it. A typical CMMS holds assets, work orders and PM schedules. It has no field for "if this fails in January we have six hours" or "the contractor who knows this unit retired, call this one instead." So that knowledge lives with the person who has been there longest.
That works until it does not:
- People leave. When a long-serving operator retires, the risk knowledge retires with them.
- Failures happen off-shift. The person on call is often not the person who knows the system best.
- Nothing gets funded. A worry in your head cannot go into a capital request. A documented scenario with a risk score can.
- Nobody prepares. The spare part, the rental agreement and the vendor number are only arranged in advance if someone wrote down that they would be needed.
Which building systems should you assess first?
Start with the systems whose failure would hurt someone, close the building or break a regulation. For most facilities that is a list of five to ten:
- Heating plant: boilers, pumps, and the controls that run them
- Cooling and refrigeration: chillers, ice plants, cold storage
- Electrical service: main switchgear, transformers, distribution
- Emergency power: generator, transfer switch, UPS
- Fire alarm and sprinkler systems
- Domestic water and backflow prevention
- Elevators, especially where they are the only accessible route
- Roof and building envelope
Then add anything with a single point of failure (one boiler, one pump, one feed) or a long lead time for replacement parts. A system you cannot repair for twelve weeks deserves a plan even if it rarely fails.
How do you run a failure scenario? The four questions
Take one system at a time. Ideally, walk to it with the person who knows it best, and answer these four questions out loud before writing anything down.
1. What happens if this system goes down?
Describe the consequences in plain terms, then check them against six kinds of impact: safety, service disruption, financial loss, regulatory, environmental and reputation. A chiller failure at an office is an uncomfortable afternoon. The same failure at an ice rink during a tournament is cancelled bookings and a refund problem. Be specific about who is affected.
2. How long can we manage without it before it becomes serious?
This is your tolerable downtime, and it is the number most teams have never put a figure on. It almost always depends on the season: a boiler outage in July can wait for parts; in January it may give you hours. Write down both. This single number tells you whether your response is "schedule a repair" or "everyone drops what they are doing."
3. What would we do, who would we call, and what would we need?
This is the response plan. Write the first three actions in order. List the vendor and their after-hours number, your contract or account number, and who in your organization can approve emergency spending or decide to close the building. Then list what the response depends on: spare parts, rental equipment, isolation valve and shutoff locations, drawings, keys and access.
4. What can we do now to reduce that risk?
This is where the exercise pays for itself. Answers usually fall into two groups. Some reduce the likelihood of failure: better preventive maintenance, a condition assessment, or a planned replacement. Others reduce the impact: stocking the parts that fail most, arranging a standby rental, installing a connection point for a temporary unit, or training a second person. Each one becomes a work order, a purchase or a line in the capital plan.
Rule of thumb: If the answer to question 2 is shorter than the time it takes to get the answer to question 3, the gap is your real risk, and question 4 is how you close it.
Worked example: the boiler plant in winter
Here is what a finished scenario looks like for a heating plant in a Canadian community facility. Note the risk score at the top: the plan sits next to it, so anyone looking at the most critical system also sees what to do about it.
Look at the last row. Four of the five actions cost little or nothing, and none of them would have happened if the scenario had stayed in someone's head. The fifth, the capital request, is now backed by a documented consequence and a tolerable downtime of six hours, which is a much easier case to make to a board than "the boiler is old."
A failure scenario template you can copy
Keep each scenario to one page. If it does not fit, it will not be read during an outage.
| Field | What to write |
|---|---|
| System and assets | Which system, and which assets are part of it? |
| Consequence | Safety, service, financial, regulatory, environmental, reputation |
| Tolerable downtime | Hours or days before it becomes serious, by season |
| First response | The first three things to do, in order |
| Who to call | Vendor, after-hours number, contract or account number |
| What you need | Spares, rentals, shutoff locations, drawings, access |
| Who decides | Who can close the building or approve emergency spend |
| Mitigation now | What reduces the likelihood or the impact this year |
| Last reviewed | Date and name |
Fill it in for your top five systems first. A rough scenario written this month is worth far more than a perfect one planned for next year.
Risk in your head vs risk you have documented
| Attribute | In someone's head | Documented |
|---|---|---|
| When the expert is away | Whoever is on call starts from scratch | The plan is on the asset, findable by anyone |
| Tolerable downtime | A feeling that it is "pretty bad" | A number, by season |
| Vendor and parts | Searched for during the outage | Listed, and spares stocked in advance |
| Capital requests | "It is old and I am worried about it" | Risk score plus a written consequence |
| Staff turnover | Knowledge leaves with the person | Knowledge stays with the facility |
| Mitigation | Good intentions | Work orders, purchases and capital items |
The same knowledge, before and after it is written down against the asset.
How often should you review failure scenarios?
At least once a year, and a review before each system's critical season is better: heating in the fall, cooling in the spring. Also review a scenario whenever one of these happens:
- The system actually fails. Compare what happened with what you wrote.
- The system is replaced, modified or expanded.
- A vendor, contract or key staff member changes.
- The use of the space changes, such as a new tenant or program.
Phone numbers and part numbers go stale faster than anything else in the plan, so check those every time.
Where AssetLab fits
When I built AssetLab, I made sure there was a place to document this information next to the asset it belongs to, and to make it easy to find and act on. The risk management module scores every asset by consequence of failure times likelihood of failure. Consequence is assessed across the same six impacts used above, and likelihood is drawn from the age, condition and maintenance history your team already records.
Alongside that score, you record the failure scenario and response plan. So the asset at the top of your risk ranking is also the one with the plan attached, and the mitigation work goes straight into work orders, PM schedules and the replacement planner. Risk scoring is part of the AssetLab 360 and Enterprise plans.
How AssetLab handles risk
Risk in AssetLab is not a separate spreadsheet. It reads the assets, inspections and work orders you already keep, and keeps the plan where the asset is.
5×5 risk matrix
Every asset scored 1 to 25 by consequence × likelihood of failure.
Six consequence dimensions
Safety, service, financial, regulatory, environmental and reputation.
Data-driven likelihood
Drawn from age, condition scores and maintenance history.
Scenarios and response plans
Document what failure means and what to do, next to the risk score.
Reusable risk profiles
Set consequence once per system and every asset in it inherits it.
From risk to action
High-risk assets feed work orders, PM schedules and the capital plan.
Frequently Asked Questions
What is a facility risk assessment?
A facility risk assessment identifies which building systems would cause the most harm if they failed, how likely that failure is, and what the team would do about it. A complete one has two halves: a risk score (consequence of failure multiplied by likelihood of failure) that ranks the systems, and a written failure scenario for each critical system that records the impact, how long you can operate without it, who to call, what you would need, and what you can do now to reduce the risk.
What questions should a failure scenario answer?
Four. What happens if this system goes down? How long can we manage without it before the consequences become serious? What would we do, who would we call, and what would we need? What can we do now to reduce that risk? The answers become the response plan, and the last one becomes a list of work orders, purchases and capital requests.
Which building systems should a new facility manager assess first?
Start with the systems whose failure would hurt people, close the building or break a regulation: heating plant and boilers, cooling and refrigeration, electrical service and switchgear, emergency power, fire alarm and sprinklers, domestic water, elevators, and the roof. Five to ten systems is enough for a first pass. Add anything with long-lead replacement parts or a single point of failure.
What is the difference between a risk score and a response plan?
A risk score tells you which systems matter most and in what order. A response plan tells you what to do when one of them fails. The score is for prioritizing maintenance and capital; the plan is for the night the boiler trips. You need both, because a ranked list with no plan still leaves the team improvising, and a plan with no ranking tends to get written for the wrong systems.
How often should failure scenarios be reviewed?
At least once a year, and also after any real failure, after a system is replaced or modified, when a key vendor or staff member changes, and before the season when the system is most critical, such as heating before winter. Phone numbers and part numbers go stale faster than anything else in the plan.
Can a CMMS store failure scenarios and response plans?
Many cannot: they hold assets, work orders and PM schedules, but have nowhere to record what failure would mean or what to do about it. AssetLab scores every asset by consequence and likelihood of failure and gives you a place to document the failure scenario and response plan alongside that score, so the plan is found where the asset is, not in a binder. Risk scoring is part of the AssetLab 360 and Enterprise plans.
Write it down before you need it
Every facility manager carries a list of systems that worry them. The difference between a worry and a plan is an hour per system and four questions: what happens, how long you can last, what you would do, and what you can do now. Have you documented your risks and response plans, or is most of that knowledge still in someone's head?
Pick your top five systems. Answer the four questions. Put the answers where the next person will find them.
If you want to see risk scoring and response plans on your own assets, we can walk you through it in about 20 minutes.
Related Articles
QR Code Asset Management: How to Tag, Track, and Maintain Assets with QR
A practical guide to QR code asset management. Learn how to implement QR tagging, mobile scanning workflows, and connect physical assets to your CMMS.
The Forest and the Trees: How Smart Asset Classification Transforms Facility Management
How AssetLab's hierarchical classification system lets you see your entire portfolio - and every individual asset - at the same time. Discover why the dual system/location hierarchy transforms capital planning, compliance tracking, and portfolio management.
Facility Compliance Tracking, and a Use We Never Designed For
Compliance in AssetLab is earned by completed work orders, not typed into a status column. How the score works, what inspectors see on a placard, and how one customer had Claude research his city requirements and create the compliance records against his asset registry.
