FRM Exam Part II · Case Study: Third-party Risk Management
Operational Resilience and Third-Party Failure Case Study
Updated 11 October 2026 · Fact-checked
Operational resilience is a firm's ability to deliver critical services through disruption. In a vendor failure case, identify the critical service, find the dependency, compare the outage with the impact tolerance, test the scenario, then judge response, communication, recovery and exit options. Answer by linking each fact to a resilience concept.
Understand Operational Resilience and Third-Party Failure Case Study
Operational resilience means a firm can prevent, adapt to, respond to, recover from and learn from disruption, so that it keeps delivering critical services (also called important business services). It differs from classic business continuity. Continuity plans protect individual systems or sites. Resilience starts from the service the customer or market relies on, and works backwards through everything that supports it.
A critical service is one whose failure would harm customers, threaten the firm's safety and soundness, or damage market stability. Examples are payments, trade settlement, and access to deposits. For each one, the firm maps the people, processes, technology, data, facilities and third parties it needs. Mapping is how you discover hidden dependencies, including fourth parties (the vendor's own suppliers).
An impact tolerance is the maximum disruption the firm will accept for a critical service, assuming the disruption happens. It is often set as a maximum time, and may also be a volume of failed transactions or customers affected. It differs from risk appetite, which is about how likely or how large a risk the firm is willing to take on. Tolerance accepts that failure will occur and asks how bad is too bad.
A vendor failure case usually follows one pattern. A cloud, payment processor or data provider has an outage or cyber incident. Several firms depend on it, which creates concentration risk. The firm checks whether its contract, monitoring, backup arrangements and exit plan worked. Then it tests whether recovery stayed inside tolerance.
Testing uses severe but plausible scenarios: vendor outage, vendor insolvency, vendor cyber attack, loss of a fourth party. Methods include tabletop exercises, failover tests, and simulations with the vendor. Lessons from tests and real incidents feed back into the framework. Outsourcing moves the activity, not the accountability: the board and senior management remain responsible.
Key formulas to remember
- Tolerance breach test
- Breach if (time to restore service) > (impact tolerance)
- Measure from the start of disruption to restoration of the critical service, not just of the vendor system.
- Recovery time against tolerance
- Time to restore = detection time + decision time + failover or recovery time
- Detection and decision delays often take up most of the tolerance. Check every component.
- Resilience mapping chain
- Critical service → people, process, technology, data, facilities → third and fourth parties
- A service is only as resilient as its weakest mapped dependency.
- Concentration indicator
- Share of critical services relying on one provider = services using provider ÷ total critical services
- A high share signals concentration risk. No universal threshold exists, so use the firm's own limits.
- Tolerance vs appetite
- Impact tolerance assumes disruption occurs; risk appetite governs likelihood and size of risk taken
- Do not treat the two as interchangeable.
How to solve Operational Resilience and Third-Party Failure Case Study questions
Use the same sequence for any vendor-failure resilience question. It keeps your answer tied to the concepts the exam tests.
- 1Identify the critical service affected and who relies on it (customers, markets, the firm).
- 2Find the third party and the dependency, including any fourth parties and shared providers.
- 3Read the impact tolerance and compare it with the actual or expected outage time.
- 4Check preventive controls: due diligence, contract terms, SLAs, monitoring, concentration limits.
- 5Assess response: detection, escalation, incident management, customer and regulator communication.
- 6Assess recovery: failover, alternative providers, manual workarounds, exit plan, testing history.
- 7Decide the best action or the main weakness and match it to one option.
- 8Rule out options that confuse tolerance with appetite, or that shift accountability to the vendor.
Quickest way: Service, tolerance, gap
When to use it: Use it for short MCQs where the stem gives an outage duration and a stated tolerance.
- Underline the critical service and the stated tolerance.
- Compute or note the outage time and compare it directly.
- If outage exceeds tolerance, look for the option that names the failed control: mapping, testing, exit or concentration.
- If the option says accountability moves to the vendor, eliminate it.
- Prefer the answer that restores the service for customers first, then fixes root cause.
Common mistakes in Operational Resilience and Third-Party Failure Case Study
Treating impact tolerance as risk appetite
Both set limits, and both are approved by senior management.
Fix: Tolerance is the maximum acceptable disruption once failure has occurred. Appetite is about the risk taken before it occurs.
Assuming outsourcing transfers responsibility to the vendor
The contract allocates the activity and sometimes liability.
Fix: The firm and its board stay accountable for the service and its resilience. Contracts cannot transfer that.
Mapping only direct vendors
Firms hold contracts only with first-tier providers.
Fix: Include fourth parties and shared infrastructure. Hidden common dependencies cause correlated failure.
Measuring outage by the vendor's system recovery
Vendors report their own restoration time.
Fix: Measure when the customer-facing critical service is actually restored, including reconciliation and backlog.
Relying on the existence of a plan instead of evidence of testing
A documented plan looks like a control.
Fix: Choose answers that cite tested scenarios, including severe but plausible vendor failure, with lessons acted on.
Picking a technical fix when the issue is governance
Candidates focus on the system that failed.
Fix: Ask whether the weakness is mapping, tolerance setting, oversight, exit planning or communication.
Worked examples
Example 1
A bank sets an impact tolerance of 4 hours for its card payments service. A processor outage begins at 09:00. Detection takes 30 minutes. Management takes 45 minutes to decide to switch to a backup provider. Failover takes 2 hours 30 minutes. Was tolerance breached, and what is the main lesson?
Show the solution
- Add the components: 30 min + 45 min + 150 min = 225 minutes.
- Convert tolerance: 4 hours = 240 minutes.
- Compare: 225 < 240, so the service was restored inside tolerance.
- Margin is only 15 minutes. Detection and decision time used 75 minutes, a third of the total.
Answer: No breach: 3 hours 45 minutes against a 4-hour tolerance. The margin is thin, so the lesson is to shorten detection and decision time and test the failover.
Example 2
A firm outsources its customer onboarding checks to one vendor. The vendor suffers a cyber incident and is offline for three days. The firm's tolerance is 24 hours. The firm had a continuity plan but never tested vendor failure, and has no alternative provider. Identify the main resilience failings.
Show the solution
- Compare outage with tolerance: 3 days = 72 hours, which exceeds 24 hours, so tolerance was breached by 48 hours.
- Testing failing: no severe but plausible vendor-failure scenario was run, so the plan was unproven.
- Exit and substitution failing: no alternative provider or manual workaround existed.
- Concentration failing: the service depended on a single vendor.
- Accountability: the firm remains responsible to customers and regulators despite outsourcing.
Answer: Tolerance was breached by 48 hours. The main failings are untested vendor-failure scenarios, no alternative provider or exit plan, and single-vendor concentration. Accountability stays with the firm.
Exam tips
- Look for the stated impact tolerance and compare it with the timeline before reading the options.
- Eliminate any option that says responsibility passes to the vendor.
- Prefer answers about tested scenarios, mapping dependencies and exit plans over generic statements about having a policy.
- Name the concept precisely: critical service, impact tolerance, fourth party, concentration risk.
- Watch for resilience versus recovery: the question may ask about the service, not the system.
Practice questions from Case Study: Third-party Risk Management
- A bank notices that several of its critical vendors all rely on the same cloud infrastructure provider. Which risk is this most directly des…
- A mid-sized bank relies on a single cloud provider to host its payments platform, core ledger and customer authentication service. The risk …
- A risk officer reviews a bank's outsourcing register and finds that the contract with a critical cloud provider lacks an exit strategy and a…
- During ongoing monitoring, a bank notes that a critical cloud provider has begun subcontracting its data-hosting to a fourth party in anothe…
- Before signing a contract with a prospective provider of loan-servicing, which activity should a bank's third-party risk framework place fir…
Operational Resilience and Third-Party Failure Case Study: frequently asked questions
What is an impact tolerance in operational resilience?
It is the maximum disruption a firm will accept for a critical service, often expressed as a time limit. It assumes the disruption happens. Firms test whether they can stay within it.
How do you test the resilience of outsourced services?
Run severe but plausible scenarios such as vendor outage, insolvency or cyber attack. Use tabletop exercises, failover tests and joint tests with the vendor. Record gaps and fix them.
How is operational resilience different from business continuity?
Business continuity focuses on recovering systems or sites. Resilience focuses on delivering critical services to customers and markets through any disruption. It also covers prevention, adaptation and learning.
Why does a third-party failure case matter for FRM Part II?
Questions are applied and case-like. You must link facts about a vendor failure to concepts such as critical services, tolerance, concentration and testing, and pick the best action.