The short answer
A residual risk score is meaningless on its own. It only says something when measured against a threshold your business agreed in advance, which is your risk appetite, and against written consequence criteria that let two different people score the same event the same way. Most organisations have a matrix and no criteria, which is why scores drift between assessors, between documents and between years. Fixing it takes one page: consequence descriptors for each discipline, a likelihood scale with a stated probability, and a rule about what happens at each band and who signs it.

Every risk register produces a verdict of some kind. Some produce a number out of 25. Plenty produce nothing but a word: high, medium, low. Either way, very few businesses can answer the question that comes next.
This one came out as a 12. Or as a Medium. Is that acceptable?
Silence, usually. Or worse, a confident yes from whoever is closest to the pressure to get the job started. A score is arithmetic and a rating is a label. Neither one is a verdict. The verdict is a decision, and if nobody made that decision in advance, it is being made informally by whoever is standing there.
Appetite, tolerance and criteria
Three words that get used interchangeably and should not be.
Risk appetite is how much risk you are willing to accept in pursuit of an objective. It is a leadership position, set deliberately, and it differs by discipline. Most businesses have a near-zero appetite for safety risk and a considerably higher one for commercial risk, which is a coherent position as long as it is stated.
Risk tolerance is what you will put up with in practice when it happens anyway. The gap between appetite and tolerance is where most organisations actually live.
Risk criteria are the written descriptors that turn both into something usable: what a level 3 consequence looks like, what “likely” means in numbers, what has to happen before work starts at each band.
Most businesses have a matrix. Far fewer have criteria. A matrix without criteria is a colour scheme.
The mistake: one consequence scale built for injuries
Open almost any risk matrix and the consequence axis reads: first aid, medical treatment, lost time, serious injury, fatality.
Then that same scale gets used to score a late delivery, a customer complaint, a defect that escaped to site, a data breach and a hydrocarbon spill. None of those are injuries. People score them by feel, the feel differs between assessors, and the register stops being comparable with itself.
You need one likelihood scale and several consequence scales, one per discipline, calibrated so that a level 4 means the same weight of bad day whichever column it sits in.
That calibration is the whole exercise, and it is where the appetite conversation actually happens. If a fatality and a $50,000 loss are both level 5 in your system, you have said something quite specific about your business, and you should have said it on purpose.
What consequence criteria look like
Here is a workable structure. Five levels, a fixed core of disciplines that apply to everyone, and a tail that changes with the customer and the contract.
It is easier to build and far easier to read if the core is split in two: harm, and business consequence. They are scored the same way and calibrated against each other, but they are different families and putting every consequence column on one line helps nobody.
| Level | Injury and harm | Environment | Information security | AI and automated decisions |
|---|---|---|---|---|
| C5 Catastrophic | Fatality or life-changing permanent disability. Prosecution likely. | Serious harm, immediate or long term, with potential for regulator prosecution. | Large-scale compromise of sensitive or personal information with serious harm likely. Regulator investigation, or an event halting operations. | Systemic incorrect, unsafe or discriminatory automated decisions affecting people’s rights. Loss of human oversight of a system with real-world effect. |
| C4 Major | Lost time injury of several days or more, with potential for permanent disability. Prosecution possible. | Off-site impact affecting neighbouring property or water. Regulator infringement notice. | Eligible data breach notified to the regulator and to affected individuals, or extended loss of a business-critical system. Customer data involved. | A pattern of wrong decisions affecting multiple people, or a system operating outside its approved purpose without anyone noticing. |
| C3 Serious | Lost time injury of a few days. | Localised impact contained within the site boundary. Reportable, with infringement potential. | Unauthorised access to or loss of personal information affecting a small number of people, or loss of a business-critical system for part of a day. | An automated or AI-assisted decision about a person is wrong and has to be reversed. Complaint or remediation required. |
| C2 Moderate | Injury requiring medical treatment, no restricted duties. | Minor impact, immediate containment, no regulator notification. | Unauthorised access to internal information that is not sensitive, or brief loss of a non-critical system. No personal information involved. | Incorrect output used internally and corrected before it reached anyone outside the business. |
| C1 Minor | Minor first aid or minor illness. | Negligible or temporary inconvenience, no notification. | Security event with no information affected, contained internally. | Incorrect output caught before use. No decision affected. |
| Level | Property and financial | Reputation | Legislative |
|---|---|---|---|
| C5 Catastrophic | Loss at or above the threshold the business has set as material. | Public or media coverage, or removal from an approved supplier list. Affects your ability to win work regardless of which customer was involved. | Externally identified and prosecuted. |
| C4 Major | Loss at the band below material. | Adverse awareness spreading beyond your customer base into the industry. Prequalification, panel membership or tender scoring affected. | Multiple externally identified infringement notices. |
| C3 Serious | Moderate loss requiring management attention. | Dissatisfaction across several customers, or a performance issue escalated to a client’s management. Puts a renewal or a referral at risk. | Externally identified breach, infringement notice received. |
| C2 Moderate | Loss absorbed within normal operating margin. | Formal complaint from one customer requiring a written response and corrective action. Relationship recoverable. | Externally identified breach, warning issued. |
| C1 Minor | Immaterial loss. | Concern raised by one customer and closed out directly. No effect on the relationship. | Internally identified potential breach, resolvable in-house. |
The variable tail. Depending on the contract, add columns for schedule (days of delay), quality (defects, failed acceptance criteria, escapes to the customer) and performance against whatever the contract measures. These are the ones to build with the customer’s own contract in front of you, because a level 4 schedule slip on a rail possession is a different animal to a level 4 on a fitout.
On the two newest columns. Information security and AI are in the core rather than the tail deliberately. Almost every business now holds personal information and runs software that makes or shapes decisions about people, and both carry statutory consequences: the notifiable data breach scheme already, and from 10 December 2026 the automated decision-making transparency requirements in APP 1.7 to 1.9. A consequence scale that cannot score a data breach or a wrong automated decision is a scale built for the work of ten years ago.
The AI column is the one most organisations have never attempted. If it feels difficult to fill in, that difficulty is the finding, and it is the same one behind assessing AI tools one at a time rather than together.
Set the financial figures against your own business, not against a template. The dollar bands are the fastest way to see whether a matrix was thought about or downloaded. If catastrophic financial loss is set at a number your business could absorb in a quarter, the scale will produce level 5 scores for events nobody in the room considers catastrophic, and people will quietly stop believing the register. This is where appetite becomes visible: the number you pick is your appetite, written down.
And do not write “major” at three different levels. If the same adjective appears at C3, C4 and C5, it is doing no work and two assessors will split them differently every time. Each level needs to fail a different test.
Rank the cells. Do not multiply them.
This is the part most matrices get wrong, and it is worth the five minutes it takes to fix.
The common approach multiplies likelihood by consequence. Five by five gives you a score out of 25, and it feels rigorous. It is not, because multiplication produces ties between events that are not equivalent.
A rare event with catastrophic consequence and a near-certain event with minor consequence can arrive at the same product. They demand completely different responses. One needs a control that stops it being possible at all. The other needs somebody to fix an irritation. A matrix that scores them identically has destroyed the information you needed.
The better approach is to rank all 25 cells from 1 to 25 and decide the order deliberately. Every cell gets a unique number, and populating the grid forces you to answer questions like: does a rare catastrophe outrank a frequent nuisance? In most businesses it should, and a ranked matrix makes you say so explicitly instead of letting the arithmetic decide by accident.
The practical payoff is that a ranked register can be sorted. With banded colours you have four piles. With 1 to 25 you have a priority order, which is what anyone allocating a budget actually needs.
If your matrix produces only words, this argument applies with more force rather than less. A register where everything is High, Medium or Low gives you three piles and no order inside them, so “Medium” ends up covering a rare catastrophe and a frequent nuisance at the same time, and the person reading it cannot tell which is which. You do not have to abandon the words. Number the cells underneath them and you get both: a rating everyone already understands, and an order you can work down.
Give likelihood a probability
The second thing that eats time in risk workshops is the argument about whether something is “possible” or “likely”. Those words mean different things to different people and there is no way to settle it.
| Level | Descriptor | Probability |
|---|---|---|
| L5 Almost certain | Common or repeating occurrence | 1 in 10 |
| L4 Likely | Known to occur, or it has happened here | 1 in 100 |
| L3 Possible | Could occur | 1 in 1,000 |
| L2 Unlikely | Not likely to occur, remote | 1 in 10,000 |
| L1 Rare | Practically impossible | 1 in 100,000 |
One in ten of what, though. A probability is only meaningful against a stated basis: per task, per shift, per year, or per exposure. Whichever you choose, write it on the matrix rather than leaving it assumed. An undefined basis is the single easiest thing for an auditor to pull on, and two assessors working to different bases will produce different scores from the same facts and never work out why.
Say what happens at each band, and who signs it
Criteria tell you what the number means. This tells you what to do about it, and it is where appetite stops being a philosophy and becomes a control.
| Band | Rating | What has to happen |
|---|---|---|
| 20 to 25 | Extreme | Work does not commence. Additional controls required to reduce the score. |
| 16 to 19 | High | Documented safe work method statement prepared by the supervisor and approved by management. Prestart meeting, confirmation of controls and responsibilities, progress monitored. |
| 9 to 15 | Medium | Documented safe work method statement prepared before work starts. Supervisor confirms the crew understands and implements the controls. |
| 1 to 8 | Low | Supervisor reviews the task with the crew at prestart. |
Two things make this table do real work.
State the target, not just the bands. Controls should be developed using the hierarchy to bring residual risk down to Low. That single sentence is your appetite, expressed in a way somebody can be held to.
Name the level of authority, and make it independent. If the person accepting a high residual risk is the same person carrying the schedule pressure, the acceptance is not an acceptance, it is a formality. Supervisor for medium, management for high, nobody for extreme, is a defensible split precisely because it moves the decision away from the pressure as the risk rises.
Reduced to what, and acceptable to whom
Worth remembering why any of this is a judgement rather than a calculation.
Skydiving with no parachute has a certain outcome. Add a parachute, a licensed operator, a tandem instructor, documented packing procedures and a reserve canopy, and the residual risk is genuinely low. Plenty of people look at that number and jump out of an aeroplane on a Saturday for fun.
Others would not. Both are looking at the same residual score.
That is what appetite is: the point at which your organisation has decided, in advance and in writing, that a number is acceptable. Without it, every score in your register is being judged against a threshold nobody wrote down, and the answer changes depending on who is asked and how busy they are.
The mechanics of how controls move a score, and why likelihood usually falls further than consequence, are covered in our article on the energy wheel and the hierarchy of controls.
Where the standards ask for this
- ISO/IEC 27001:2022, clause 6.1.2. The strongest hook of the lot. It requires you to define and apply an information security risk assessment process that establishes and maintains information security risk criteria, including risk acceptance criteria. If you hold 27001 and cannot produce written acceptance criteria, that is a nonconformity rather than a nice-to-have.
- ISO/IEC 42001:2023, clauses 6.1.2 and 6.1.4. AI risk assessment and the AI system impact assessment. The structure mirrors 27001: the process is established in clause 6 and performed under clauses 8.2 and 8.4. The impact assessment asks what a system does to the people it affects, which is a consequence question and needs a consequence scale to answer it.
- ISO 45001:2018, clause 6.1.2.2. Assessment of OH&S risks and other risks to the OH&S management system. Hazard identification sits separately at 6.1.2.1, which is the hazard-versus-risk split expressed as two sub-clauses.
- ISO 14001:2015, clause 6.1.2. Environmental aspects, and determining which are significant. That significance test is a consequence criterion in all but name, and it is the clause most likely to embarrass a business that has never written criteria down.
- ISO 9001, clause 6.1. Risks and opportunities, and deliberately undemanding. It prescribes no method and no matrix, which is why the quality column is almost always the weakest one on the page.
- ISO 31000 supplies the vocabulary, including risk criteria and risk appetite.
If you run an integrated system, one set of criteria should serve all of them. Running separate matrices for safety, environment, quality and information security is how you end up scoring the same event four different ways in four different documents, and it is the reason ISO 42001 and ISO 27001 share so much machinery once the risk framework underneath is common.
A test for your next management review
Fifteen minutes, and it produces a real finding.
Take three closed risks from your register. Hand them to two people separately and ask each to score likelihood and consequence from the criteria. Then compare.
If the scores differ, the problem is the criteria, not the people. That is the finding, it is objective, and it is far more useful than another review meeting where everyone agrees the register looks fine.
The free template
The obvious objection to all of the above is that this article has just told you not to download a matrix, and is now offering you one. Fair. So, to be clear about what it is: a structure with a worked example inside it, not a set of numbers to adopt. Every cell you are meant to replace is shaded yellow and written in italics, and the financial bands are the first thing to change.
Six sheets. A 5×5 matrix with all 25 cells ranked rather than multiplied. Consequence criteria across ten columns, information security and AI among them. A likelihood scale with probabilities and a stated basis. The ISO 45001 clause 8.1.2 hierarchy, with a note against each item on whether it moves likelihood or consequence. An action table naming who may accept each band. And a sheet explaining why each of those choices was made, so you can disagree with them deliberately.
It is a spreadsheet, it is free, and there is nothing to fill in to get it.
Download the risk matrix template (XLSX)
Start with the consequence sheet. If the financial bands still say what we wrote, nobody has had the appetite conversation yet. If you want a sheet by sheet walk through of the file, including a worked example scored end to end, that is in our guide to the risk matrix template.
Frequently asked questions
What is the difference between risk appetite and risk tolerance?
Appetite is how much risk you are willing to take in pursuit of an objective, set deliberately in advance. Tolerance is what you will put up with when a risk materialises. A business can have a low appetite for safety risk and still tolerate a good deal of it in practice, and the gap between the two is usually where the audit findings come from.
Do we need separate consequence scales for quality, safety and environment?
Separate scales, one shared likelihood scale, and all calibrated against each other. A single consequence axis written for injuries cannot meaningfully score a defect, a spill, a data breach or a wrong automated decision, so people score by feel and the register stops being comparable.
Should information security and AI have their own consequence columns?
Yes, and in the core rather than as an optional extra. A data breach has a statutory consequence path of its own under the notifiable data breach scheme, and from 10 December 2026 automated decisions carry disclosure obligations under APP 1.7 to 1.9. ISO 27001 clause 6.1.2 goes further than the other standards and requires written risk acceptance criteria in terms. If you hold 27001 or are heading toward ISO 42001, this is the column an auditor is most likely to ask to see.
Who should approve a high residual risk?
Somebody far enough from the delivery pressure that the approval means something. Write the authority level against each band, and set at least one band at which the answer is that work does not proceed.
Can we use one matrix across all our standards?
Yes, and you should. One likelihood scale, one set of bands, one action table, and a consequence column per discipline. The alternative is several matrices that disagree, which an auditor will find by opening two documents.
Where Streamline fits
Most of the risk frameworks we are asked to look at are not wrong. They are inherited. Somebody built the matrix once, it has been copied into every document since, and nobody has tested whether two people using it reach the same answer.
Streamline builds and reviews risk frameworks as part of integrated management system design for Australian organisations, across ISO 9001, ISO 45001, ISO 14001, ISO 27001 and ISO 42001. You deal directly with a practising ISO Lead Auditor, so the criteria you end up with are the ones that hold up when they are tested rather than the ones that look tidy. A gap analysis is the usual starting point, and ISO mentoring is the option where you would rather your own people did the building.
Run the two-assessor test first. If the scores match, you are in better shape than most. Get in touch if they do not.
Stay in the Loop
Get an email when we post an article. Your email address will not be used for marketing, and you can unsubscribe at any time.
We handle your details in line with our privacy policy.











