FRM Exam Part II · Risk Reporting
Operational Risk Reporting and Resilience Metrics Explained
Updated 11 October 2026 · Fact-checked
Operational risk reporting gives management, the board and regulators a clear view of loss events, key risk indicators (KRIs), control health and resilience. Good reports are timely, accurate and decision-focused. To solve questions, identify the audience, match the metric to the risk, compare it to appetite and thresholds, and name the action required.
Understand Operational Risk Reporting and Resilience Metrics
Operational risk is the risk of loss from inadequate or failed processes, people, systems or external events. Reporting turns raw information about these risks into something a decision-maker can act on. A report that cannot trigger a decision has failed, however detailed it is.
There are four main building blocks. Loss event data records what went wrong: date of occurrence, date of discovery, date of accounting, gross loss, recoveries, net loss, business line, event type and cause. Key risk indicators (KRIs) are forward-looking metrics that signal rising risk before losses occur, such as staff turnover, system outages, failed trades or overdue audit actions. Control indicators show whether controls work. Resilience metrics show whether critical services can keep running through disruption.
A KRI differs from a KPI. A KPI measures how well a process is performing against a business goal, such as transactions processed per hour. A KRI measures the exposure to risk that could stop the goal being met, such as the percentage of transactions needing manual repair. Some metrics can serve as both, but the purpose differs: KPI is about performance, KRI is about risk. KRIs are only useful with thresholds. Typically these are set in tiers, often green, amber and red, linked to risk appetite and tolerance. A breach triggers escalation to a named owner.
Operational resilience is the ability to deliver critical operations through disruption. Firms identify critical services, set impact tolerances (the maximum tolerable disruption, such as a maximum outage duration), map the people, technology, third parties and data that support them, and test against severe but plausible scenarios. Resilience reporting then tracks whether tested recovery times sit within tolerance, how many scenarios failed, and which dependencies are weak.
Reports go to different audiences. The board needs a short, aggregated view tied to risk appetite and the top emerging risks. Senior management needs trends, root causes and action status. Business lines need granular detail. Regulators need accurate, consistent data. Basel principles on risk data aggregation and reporting (BCBS 239) expect reports to be accurate, complete, timely and adaptable, and to be based on strong data governance.
Key formulas to remember
- Net loss
- Net loss = Gross loss − Recoveries
- Recoveries include insurance and other amounts recovered. Report gross, recovery and net amounts separately.
- KRI threshold logic
- Green: within appetite | Amber: approaching limit | Red: limit breached → escalate
- Thresholds are set by the firm. The principle is that a breach triggers a defined escalation and action.
- Impact tolerance test
- Tested recovery time ≤ impact tolerance → within tolerance
- If the tested recovery time is longer than the tolerance, the service is outside tolerance and remediation is needed.
- Loss event date types
- Date of occurrence ≤ Date of discovery ≤ Date of accounting
- Basel loss data collection records all three. Use them to analyse detection lags.
- Event reporting rate
- Rate = Number of events ÷ Volume of activity
- Normalising by volume (for example per 10,000 transactions) lets you compare periods and units fairly.
How to solve Operational Risk Reporting and Resilience Metrics questions
Use this sequence for any question on operational risk reporting, KRIs, loss data or resilience metrics.
- 1Identify the audience (board, management, business line or regulator) and what decision they need to make.
- 2Classify the metric: loss data (backward-looking), KRI (forward-looking), control indicator, KPI or resilience measure.
- 3Check data quality: completeness, accuracy, consistency of event types and timing of dates.
- 4Compare the metric with its threshold, risk appetite or impact tolerance and note the status.
- 5Look at trend and root cause, not just the current value. Normalise by volume if needed.
- 6Decide the escalation or action and the owner, with a deadline.
- 7Pick the answer that is timely, aggregated for the audience and linked to risk appetite, rejecting options that only add detail.
Quickest way: Audience, type, threshold
When to use it: Use when you have about 90 seconds for a scenario-based multiple-choice question.
- Ask who reads the report: board gets summary and appetite, regulator gets consistent data.
- Ask whether the metric looks back (losses) or ahead (KRI).
- Check if a threshold or tolerance is breached and whether escalation happened.
- Eliminate options that report only losses, lack thresholds or skip an owner.
- Choose the option that is forward-looking, actionable and tied to appetite or tolerance.
Common mistakes in Operational Risk Reporting and Resilience Metrics
Treating KRIs and KPIs as the same thing.
Both are numbers on a dashboard and some overlap.
Fix: Ask whether the metric measures performance toward a goal (KPI) or exposure to a risk that could cause a loss (KRI).
Calling loss data a leading indicator.
Losses feel like risk information.
Fix: Loss data is backward-looking. KRIs are the forward-looking signal. Loss data validates and calibrates KRIs.
Choosing a KRI with no threshold or owner.
Students focus on picking the metric, not using it.
Fix: A KRI needs thresholds linked to appetite, escalation steps and an accountable owner.
Reporting only net loss.
Net looks like the real cost.
Fix: Record gross loss, recoveries and net loss separately. Gross shows the true exposure; recoveries may be uncertain or delayed.
Confusing impact tolerance with risk appetite.
Both set limits.
Fix: Risk appetite is the level of risk a firm accepts to pursue objectives. Impact tolerance is the maximum tolerable disruption to a critical service, assuming disruption occurs.
Giving the board maximum detail.
More data feels safer.
Fix: Boards need aggregated, decision-focused reports with trends, breaches and emerging risks. Detail sits lower in the chain.
Worked examples
Example 1
A bank's payments unit processed 2,50,000 transactions in March and 3,00,000 in April. Manual repairs were 1,250 in March and 1,800 in April. The amber threshold is 0.60% and the red threshold is 0.80%. What is the April status, and what does it tell management?
Show the solution
- March rate = 1,250 ÷ 2,50,000 = 0.50%.
- April rate = 1,800 ÷ 3,00,000 = 0.60%.
- Compare April with thresholds: 0.60% reaches the amber level of 0.60% but is below red at 0.80%.
- Trend: the rate rose from 0.50% to 0.60% even though volume rose, so the increase is not just volume.
- Interpretation: this KRI is forward-looking, since high manual repair rates signal weaker straight-through processing and higher error and loss risk.
Answer: April is at amber (0.60%), up from 0.50%. Management should investigate the root cause and assign an owner before the KRI reaches red.
Example 2
A critical payments service has an impact tolerance of 4 hours maximum outage. A severe but plausible scenario test shows recovery took 6 hours. What should the resilience report to the board say?
Show the solution
- Compare tested recovery with tolerance: 6 hours > 4 hours.
- The service is outside impact tolerance by 2 hours.
- This is a resilience gap, even if no real incident has occurred, because the test shows the firm could not stay within tolerance.
- The report should identify the cause (for example a dependency on a third-party provider or a manual recovery step), propose remediation with an owner and timeline, and say when retesting will occur.
Answer: The report should state that the service failed its tolerance test by 2 hours (6 hours against 4 hours), explain the cause, and set out remediation, an owner and a retest date.
Exam tips
- Match the metric to its role: loss data looks back, KRIs look ahead, and KPIs measure performance.
- When a question mentions the board, choose the aggregated, appetite-linked report over the detailed one.
- For resilience questions, remember the test is against impact tolerance for critical services, not against average performance.
- Link data quality answers to BCBS 239 ideas: accuracy, completeness, timeliness and adaptability.
- Watch for volume effects: a rising count may be a flat rate. Normalise before judging.
Practice questions from Risk Reporting
- A bank relies on manual spreadsheets to aggregate credit, market and operational risk data from several legal entities. An internal review f…
- A bank's internal audit reviews its BCBS 239 compliance. It notes: (1) reports reconcile to the general ledger, (2) the bank uses extensive …
- A bank's operational risk team distributes its monthly risk report by email to a broad list that has grown over several years. An internal r…
- An operational risk manager wants to confirm that the quarterly risk report to senior management is timely and reliable. Which control best …
- A bank sets a reporting escalation protocol: any single operational loss event above USD 500,000 must reach the CRO within 24 hours; the mon…
Operational Risk Reporting and Resilience Metrics in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Operational Risk Reporting and Resilience Metrics: frequently asked questions
What is the difference between KRI and KPI in operational risk?
A KPI measures how well a process performs against a business goal. A KRI measures exposure to a risk that could cause loss or disruption. KRIs need thresholds and escalation, and they are meant to warn early.
How should operational risk loss events be reported?
Record each event with dates of occurrence, discovery and accounting, gross loss, recoveries, net loss, event type, business line and cause. Report aggregated trends to management and the board, and provide consistent data to regulators.
What are operational resilience metrics?
They show whether critical services can continue through disruption. Examples include tested recovery time against impact tolerance, number of failed scenario tests, and the status of third-party dependencies.
Are KRIs backward-looking or forward-looking?
KRIs are designed to be forward-looking. They signal rising risk before losses appear. Loss data is backward-looking and is often used to check whether KRIs are predictive.