Feature Operational Risk
Operational Resilience vs. Business Continuity: Why Your BCP Isn't Enough — and How to Close the Gap
Business continuity planning is about recovering from disruptions. Operational resilience is about ensuring you never exceed the maximum tolerable impact on critical services — before, during, and after disruptions. US regulators are converging on the operational resilience standard. Here's what it means for your program.
Table of Contents
TL;DR
- Business continuity planning answers: “How do we recover from this disruption?” Operational resilience answers: “How do we ensure our most important services never breach their maximum tolerable disruption?” These are related but different questions — and US regulators are increasingly expecting both
- The UK PRA required full operational resilience compliance from March 31, 2025; DORA has applied across EU financial entities since January 17, 2025; the US interagency framework (SR 20-24) applies to large banks now and is signaling direction for all sizes
- The five building blocks: identify your important business services, map their end-to-end dependencies, set impact tolerances, test under severe but plausible scenarios, document communications plans
- The most common US gap: testing. Most institutions have written impact tolerances that have never been stress-tested against a realistic disruption scenario
On July 19, 2024, a defective content update from CrowdStrike Falcon crashed an estimated 8.5 million Windows devices worldwide in hours. Banks, payment processors, and trading platforms experienced simultaneous outages. Most had BCPs. Most had RTOs. Most had documented recovery procedures for system failures.
What most didn’t have: a tested answer to the question “if our payment processing system is down for 12 hours, can we continue accepting customer funds and settling transactions through any other pathway — and at what point does that disruption breach our stated service commitment to customers and counterparties?”
That’s the difference between business continuity planning and operational resilience. One asks how you recover from a specific event. The other asks how you ensure your critical service delivery stays within a defined boundary of disruption — before, during, and after events that BCPs might not have anticipated.
Global regulators figured this out first. US regulators are catching up. Here’s what your program needs to account for.
The Fundamental Distinction
Traditional BCP design is event-centric: disaster → activation → recovery → return to normal. It’s organized around the disrupting event and the systems or processes directly affected.
Operational resilience is outcome-centric: identify your most important services → define the maximum disruption those services can tolerate → map every dependency behind them → verify you can stay within tolerances under realistic severe scenarios.
| Dimension | Business Continuity Planning | Operational Resilience |
|---|---|---|
| Primary unit of analysis | System or process | Business service (customer outcome) |
| Core question | How do we recover? | How do we stay within tolerable limits? |
| Trigger for planning | Specific disruptive events | Any scenario affecting service delivery |
| Key metric | RTO/RPO | Impact tolerance (time + magnitude) |
| Third-party scope | Contracted vendors | Full dependency chain including nth-party |
| Testing model | Recovery drills | Scenario testing against tolerances |
| Regulatory expectation (UK/EU) | Required (but not sufficient) | Required |
The BCP doesn’t disappear in an operational resilience framework — it becomes a component. Recovery procedures, crisis communications, and alternate processing arrangements all feed into operational resilience. The difference is the organizing principle: service continuity outcomes, not event-by-event recovery.
Where Global Standards Are
The international framework has been converging for several years:
UK PRA SS1/21 (Supervisory Statement 1/21) established the foundational framework: financial institutions must identify their important business services, set impact tolerances for each, map dependencies end-to-end, and test whether they can operate within tolerances by March 31, 2022 — with full compliance including successful tolerance testing required by March 31, 2025. PRA-regulated firms spent four years building and testing these frameworks. That body of practice is now informing what “good” looks like for every global institution.
DORA (EU Digital Operational Resilience Act) took full effect January 17, 2025, applying across 21 types of EU financial entities — banks, insurers, investment firms, payment institutions, crypto-asset service providers, and others — plus their critical ICT third-party service providers. DORA’s focus is specifically ICT operational resilience: digital risk management, incident classification and reporting, DORA-required testing (TLPT for significant entities), and ICT third-party risk management. The advanced TLPT testing requirements under DORA Article 26 are detailed here.
Basel Committee “Principles for Operational Resilience” (March 2021) established the multilateral standard: seven principles covering governance, operational risk management, business continuity planning, mapping of interconnections, third-party dependency management, incident management, and resilient ICT infrastructure. The US agencies endorsed these principles as supervisory direction.
Where US Regulation Currently Sits
The formal US framework for operational resilience — SR 20-24 / OCC Bulletin 2020-94, the interagency “Sound Practices to Strengthen Operational Resilience” — was issued November 2, 2020 by the OCC, Federal Reserve, and FDIC.
The critical scoping point: SR 20-24 formally applies to domestic banking organizations with average total consolidated assets of $250 billion or more, or $100 billion or more with certain risk factors. Community banks and most fintechs are not directly in scope of SR 20-24.
But the framework shapes the examination environment for everyone — for two reasons.
First, large bank practices become the industry benchmark. When examiners at $5B community banks have spent years examining $500B institutions against operational resilience standards, those expectations don’t disappear when they walk into a smaller institution’s examination. The questions evolve.
Second, the FFIEC Business Continuity Management Booklet — which applies broadly — incorporates resilience concepts. Testing requirements, dependency mapping expectations, and service continuity documentation are already in the FFIEC framework that every institution examiner is trained on.
The current examiner posture for non-SR 20-24 institutions: you won’t be asked whether you’ve set formal impact tolerances with the UK PRA’s precision. But you will be asked whether you’ve identified your critical operations, tested your recovery procedures under realistic scenarios, and verified that your critical vendors can actually deliver on their BCPs. What FFIEC and OCC 2023-17 require for third-party BCP verification is detailed here.
The Five Building Blocks
Whether you’re building a formal operational resilience framework or extending your existing BCP to incorporate resilience concepts, five elements define the structure:
1. Identify Your Important Business Services
The first step is defining which services you provide to external parties — customers, counterparties, the market — that would cause material harm if disrupted beyond a defined threshold.
This is different from listing your critical systems. A payment processing system is a system; “processing customer deposit withdrawals” is a business service. The distinction matters because service continuity can sometimes be maintained even when a specific system is down — if you’ve mapped the dependencies and built the alternatives.
For a community bank: accepting deposits, processing withdrawals, originating loans, and maintaining customer account access are typical important business services. For a fintech: payment processing, account funding, transaction settlement, and customer authentication are typical candidates.
2. Map Dependencies End-to-End
For each important business service, map every resource that the service depends on: people (key individuals, skill sets), processes (manual and automated), technology (systems, networks, data), data (where it lives, how it flows), and third-party providers (who does what, who their subcontractors are).
This mapping is where most institutions discover gaps they didn’t know existed. A payment service may depend on a core processor, which depends on a cloud provider, which has its own concentration risks. A fraud detection function may depend on a third-party AI vendor whose training data update schedule affects detection performance in ways that nobody at the bank has traced.
The October 2025 AWS outage reinforced this dynamic: financial institutions that hadn’t mapped their critical service dependencies to specific AWS regions discovered the exposure in real-time rather than on paper.
3. Set Impact Tolerances
An impact tolerance defines the maximum level of disruption to each important business service that your institution can sustain without causing unacceptable harm. It has two dimensions: time and magnitude.
Time: how long the disruption can last before causing harm (“customer withdrawals must be available within 4 hours of any system failure”).
Magnitude: how severe the disruption can be (“no more than 500 customers unable to access their accounts at any point during an outage”).
Impact tolerances differ from RTOs in an important way: RTOs are internally defined technical targets. Impact tolerances are calibrated to external consequences — customer harm, counterparty impact, systemic risk, reputational damage. You may have an RTO of 2 hours for your core banking system and simultaneously realize that customers are impacted within 30 minutes of that system going down. The impact tolerance forces you to reconcile those two facts.
4. Test Against Severe But Plausible Scenarios
This is the building block where US institutions most consistently fall short — and where UK PRA enforcement observations repeatedly identified gaps.
Testing your impact tolerances means selecting scenarios severe enough to actually stress the service — not tabletop exercises that confirm you have a recovery plan, but scenarios designed to find where your tolerance would be breached. Useful scenario types:
| Scenario Category | Example for Financial Institutions |
|---|---|
| Technology disruption | Primary data center failure + backup unavailable |
| Cyber event | Ransomware locks core banking and backup systems simultaneously |
| Third-party failure | Core processor or cloud provider extended outage |
| Key-person concentration | CFO, head of operations, and CTO simultaneously unavailable |
| Communications failure | Primary and secondary telecom provider simultaneous outage |
| Liquidity event | Deposit run + collateral call + funding market closure |
The test should answer specifically: for this scenario, at what point does our service delivery breach the impact tolerance? What is the exact failure point, and have we addressed it?
5. Document Your Communications Framework
Operational resilience requires documented communications plans that specify how the institution communicates with customers, counterparties, regulators, and staff during a disruption — across scenarios where primary channels may themselves be unavailable.
This is not just a PR or customer service question. Regulatory notification requirements (OCC incident reporting, SEC 4-business-day rule, DORA notification timetables) impose specific timelines. Your communications framework must be able to operate even when the systems that normally produce those notifications are the ones that are down.
Where Most US Institutions Actually Are
Honest assessment of where the industry sits in 2026:
What most institutions have: Written BCPs, documented RTOs and RPOs, vendor BCPs on file, crisis communication plans, and periodic tabletop exercises.
What most institutions don’t have: Formally defined important business services as the unit of analysis, tested impact tolerances (written tolerances are common; tolerances verified through scenario testing are not), end-to-end dependency maps that include third-party and fourth-party chains, and scenario testing severe enough to find the actual failure point rather than confirm existing procedures.
The practical consequence: when an examiner asks “at what point would your payment processing service exceed your maximum tolerable disruption?” — most institutions can answer with a system RTO. They can’t answer with a tested, evidence-backed tolerance that accounts for the full dependency chain and validates that the recovery pathway actually works.
Building Toward Operational Resilience Without Rebuilding From Zero
For institutions that aren’t starting with a blank page, the path from BCP to operational resilience is additive:
Extend your critical process inventory into a business service map. Your existing BCP likely has a list of critical processes. Reframe those processes as business services and identify which ones have external customer or counterparty dependencies. That reframing changes how you scope the exercise.
Add impact tolerances to your existing RTOs. For each important business service, ask: beyond “when can the system recover?” — “at what point does service disruption cause harm, and what is the maximum we can tolerate?” Document the answer with both time and magnitude dimensions.
Test one scenario hard. Rather than building 12 scenarios simultaneously, pick the one most likely to actually breach your tolerances — usually a simultaneous technology and vendor failure — and test it rigorously. Document where the tolerance is breached. Address that gap. That’s more valuable than 10 well-behaved tabletops.
Map your top three dependencies for each critical service. Full end-to-end mapping is valuable but time-consuming. Start with the three or four vendor or system dependencies most likely to cause service disruption and trace them one level deeper than you normally do.
So What?
Examiners won’t formally evaluate community banks and mid-size fintechs against SR 20-24’s standards tomorrow. But the questions are already showing up at smaller institutions — because the framework that examiners are trained on, the events that drive supervisory concern (CrowdStrike, Change Healthcare, SVB), and the regulatory signals from DORA and UK PRA are all pointing in the same direction.
The most defensible position isn’t to wait for a formal US operational resilience rulemaking at your asset tier. It’s to extend your existing BCP in the direction regulators are already headed: identify your important services, set tolerances with external-impact calibration, and test one realistic scenario before your next examination cycle.
Your BCP is a foundation. Operational resilience is where examiners are looking to see what you built on it.
The Business Continuity & Disaster Recovery Kit includes critical service identification templates, BIA methodology, vendor BCP verification checklists, and scenario testing frameworks — designed to bring your program to FFIEC BCM standards and lay the groundwork for the operational resilience documentation your examiners are increasingly expecting.
Related Reading
- Third-Party Dependent BCP: What FFIEC BCM and OCC 2023-17 Require You to Verify About Vendor Continuity
- Communications Failure BCP: When Your Telecom, Email, and Collaboration Platforms Go Down Simultaneously
- DORA Article 26 TLPT: Who Gets Designated for Threat-Led Penetration Testing in 2026 and How to Prepare Before Your NCA Calls
Sources:
- SR 20-24: Interagency Paper on Sound Practices to Strengthen Operational Resilience — Federal Reserve (November 2020)
- OCC Bulletin 2020-94: Operational Risk — Sound Practices to Strengthen Operational Resilience
- What Is the Digital Operational Resilience Act (DORA)? — IBM Think
- Digital Operational Resilience: A Compliance Priority for 2025 — Steptoe (January 2025)
- UK Operational Resilience Rules: Are You Ready for 31 March 2025? — Sidley Austin
- Closing 2025 and Reframing Resilience for 2026 — GRC 20/20 Research
- A Guide to Operational Resilience for Financial Institutions — Ncontracts
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What's the practical difference between a BCP and an operational resilience framework?
Does SR 20-24 apply to community banks and mid-size fintechs, or just large institutions?
What is an 'impact tolerance' and how is it different from an RTO?
How is DORA's approach to operational resilience different from the UK PRA framework?
Where do most US banks fall short on operational resilience today?
Do we need a separate operational resilience team, or can our BCP team own this?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Keep reading
Related posts.
Operational Risk
Risk Assessment Template in Excel: Build the Evidence Trail, Not Just the Heat Map
Build a risk assessment template in Excel that preserves evidence, challenge, approvals, and score history—not just a polished heat map.
Jul 23, 2026
Operational Risk
FedNow's Network Intelligence API Launched in April 2026. Your Fraud Risk Program Probably Hasn't Caught Up.
On April 28, 2026, the Federal Reserve made pre-payment network-level fraud intelligence available to every FedNow participant. The data — receiver account behavioral trends derived from system-wide FedNow activity — is available before a transaction is approved. Most institutions haven't updated their fraud policies, controls, or KRIs to account for what this changes.
Jul 21, 2026
Operational Risk
3,383 Incidents Later: What DORA's First ICT Data Reveals About Your Operational Risk Program
The ESAs published their first DORA ICT incident report in June 2026 — 3,383 major incidents, nearly one-third from third-party failures, only 10% cyber-related. Here's what the data means for your operational risk program.
Jul 16, 2026