FRM Exam Part II · Introduction to Operational Risk and Resilience
Operational Resilience Principles for FRM Part II
Updated 11 October 2026 · Fact-checked
Operational resilience is a firm's ability to deliver its critical operations through disruption, within set impact tolerances. You identify critical operations, set the maximum tolerable disruption, map the people, processes, technology and third parties behind them, then test with severe but plausible scenarios. Assume disruption will happen.
Understand Operational Resilience Principles
Traditional operational risk asks: how do we stop bad events and measure losses? Operational resilience asks a different question: when a disruption happens anyway, can we keep serving customers and the market? It shifts focus from preventing every failure to limiting the harm of failure.
The regulatory view, as in the Basel Committee's operational resilience principles, starts with critical operations. These are the activities whose failure would harm customers, the firm's safety and soundness, or financial stability. You decide what is critical by their impact on others, not only by their revenue.
Next comes the impact tolerance. This is the maximum level of disruption you will accept for a critical operation, usually expressed as a time limit, a volume of affected transactions or a degree of customer harm. Regulators expect you to set it assuming the disruption has already happened. It is different from risk appetite, which is the level of risk you are willing to take before events occur and which is about likelihood and loss. Impact tolerance is about how much damage you can bear after the event.
To meet tolerance you map dependencies: the people, processes, technology, data, facilities and third parties behind each critical operation. Mapping exposes single points of failure and concentration in outsourced providers. Business continuity plans and recovery arrangements then restore services within tolerance. Finally, you test using severe but plausible scenarios, such as a cloud outage, a cyberattack or loss of a key site. A scenario must be extreme enough to stretch you but realistic enough to happen. Test results should show whether you can stay within tolerance, and gaps feed back to the board.
Governance matters too. The board and senior management approve the approach, set tolerances and oversee remediation. Resilience builds on, and links to, operational risk management, business continuity, third-party risk, cyber risk and incident management.
Key formulas to remember
- Operational resilience definition
- Resilience = deliver critical operations through disruption, within impact tolerance
- The test is delivery of the service, not avoidance of the event.
- Impact tolerance vs risk appetite
- Impact tolerance = maximum tolerable harm after disruption; Risk appetite = risk accepted before events
- Tolerance assumes the failure has happened. Appetite looks at likelihood and loss beforehand.
- Resilience cycle
- Identify critical operations → set tolerances → map dependencies → continuity and controls → test with scenarios → remediate
- Good for ordering questions about the sequence.
- Recovery time test
- Actual recovery time ≤ impact tolerance time
- If recovery exceeds tolerance, the firm has a resilience gap to fix.
How to solve Operational Resilience Principles questions
Use this method for scenario questions on resilience.
- 1Identify what is being asked: a definition, a distinction, a sequence step or a judgement on a case.
- 2Find the critical operation and who would be harmed if it failed: customers, the firm or the market.
- 3Check whether the stem is about before-event likelihood (risk appetite, prevention) or after-event harm (impact tolerance, recovery).
- 4Look for dependencies: technology, people, data, facilities and third parties, and any single point of failure or concentration.
- 5Compare recovery capability with the stated tolerance, for example hours needed against hours allowed.
- 6Check the scenario is severe but plausible and tests the end-to-end service, not one system.
- 7Choose the option that matches the regulatory principle and includes board oversight and remediation of gaps.
Quickest way: Tolerance, map, test
When to use it: When time is short and options look similar.
- Ask: is this about preventing the event or surviving it? Surviving means resilience.
- If a number appears, compare recovery time with tolerance.
- Prefer answers that focus on critical operations and end-to-end mapping, not on the whole firm or a single system.
- Reject answers that tie tolerance to expected loss or capital.
Common mistakes in Operational Resilience Principles
Treating impact tolerance as the same as risk appetite
Both set limits and both use board approval.
Fix: Appetite is about risk taken before events. Tolerance is the maximum disruption accepted once an event has happened.
Defining critical operations by revenue only
Firms naturally think about profit.
Fix: Define them by the harm their failure causes to customers, the firm's safety and soundness, or financial stability.
Assuming resilience means preventing all disruption
Operational risk training stresses controls and prevention.
Fix: Resilience assumes disruption occurs. The goal is to keep delivering within tolerance.
Mapping only internal systems
Third parties feel outside the firm's control.
Fix: Map all dependencies, including outsourced providers, cloud services, data and people, and watch for concentration.
Using mild or extreme-impossible scenarios
Candidates forget the phrase severe but plausible.
Fix: The scenario must be tough enough to test tolerance yet realistic enough to occur.
Equating business continuity with resilience
Continuity plans are the most familiar tool.
Fix: Continuity is one component. Resilience also covers tolerances, mapping, testing, governance and learning.
Worked examples
Example 1
A bank's payments service is a critical operation. The board sets an impact tolerance of 4 hours of outage. A scenario test of a data centre failure shows recovery takes 9 hours. What should the bank conclude and do?
Show the solution
- Compare recovery time with tolerance: 9 hours against 4 hours.
- 9 hours exceeds 4 hours by 5 hours, so the bank cannot stay within tolerance in this scenario.
- This is a resilience gap, even though the scenario is severe but plausible.
- The bank should map the dependencies causing the delay, such as a single data centre or provider, and fix them.
- Remediation could include alternative sites or faster failover, followed by retesting and reporting to the board.
Answer: The bank breaches its tolerance by 5 hours. It must investigate dependencies, remediate and retest, and report to the board.
Example 2
A risk manager says: 'Our impact tolerance for card authorisation is the same as our operational risk appetite, so we need only one metric.' Is this correct?
Show the solution
- Risk appetite describes the level of risk the firm accepts before events, often framed by likelihood and loss.
- Impact tolerance describes the maximum disruption to a critical operation the firm can accept once it has occurred.
- The two are linked but answer different questions, and tolerance is set assuming the failure has happened.
- Tolerance is expressed in terms such as maximum outage time or customers affected, which differ from loss-based appetite metrics.
- So one metric cannot replace both.
Answer: Incorrect. Appetite is pre-event risk acceptance. Tolerance is post-event maximum disruption, and the firm needs both.
Exam tips
- Look for the keywords after the event or assume disruption has occurred. They signal impact tolerance.
- Expect questions that ask you to separate operational risk management from operational resilience.
- When recovery time and tolerance are given, a simple comparison usually decides the answer.
- Third-party concentration and single points of failure often appear as the hidden gap in case stems.
- Choose answers that include board oversight and testing with severe but plausible scenarios.
Practice questions from Introduction to Operational Risk and Resilience
- A trader at a bank deliberately books fictitious hedging trades to hide losses on his own positions, and no outside party is involved. Under…
- A bank records four events in its loss database. Event 1: a clerk's keying error causes a USD 200,000 payment to the wrong party, recovered …
- A trading desk mistakenly enters a trade as a sell instead of a buy, and the error is discovered after settlement, causing a loss. Which Bas…
- A custodian bank's operations clerk keys the wrong settlement date on a client's securities transfer instruction. The mistake causes a faile…
- A bank's BIC is EUR 200 million. Its ILM is 1.2. Operational risk RWA is derived from the minimum capital requirement. What are the capital …
Operational Resilience Principles: frequently asked questions
What is the difference between operational risk and operational resilience?
Operational risk focuses on identifying, preventing and measuring losses from failed processes, people, systems or external events. Operational resilience assumes some disruption will occur and focuses on keeping critical operations running within tolerance. The two complement each other.
What is impact tolerance?
It is the maximum level of disruption to a critical operation that a firm is prepared to accept, often stated as a time limit or customer harm. It is set on the assumption that the disruption has already happened.
How is impact tolerance different from risk appetite?
Risk appetite is the risk a firm accepts before events happen. Impact tolerance is the most disruption it can bear after an event. Appetite looks at likelihood and loss, while tolerance looks at harm to the service.
Why does mapping dependencies matter?
Mapping shows the people, processes, technology, data, facilities and third parties that a critical operation relies on. It reveals single points of failure and concentration, so you know what to protect and test.