Feature Business Continuity
The October 2025 AWS Outage Was a BCP Exam That Most Fintechs Didn't Know They Were Taking
A 15-hour AWS outage in October 2025 locked customers out of financial accounts and froze transactions across 1,000+ companies — and exposed how few fintechs had actually stress-tested their cloud concentration risk. Here's what the FFIEC BCM handbook requires, what the OCC's 2026 report found, and what your BCP needs to say about single-provider dependency.
Table of Contents
On October 20, 2025, a malfunction in an internal network load balancer monitoring subsystem at an AWS data center in northern Virginia cascaded into a 15-hour outage that affected over a thousand companies, generated more than six and a half million user reports, and cost the global economy more than a billion dollars. Financial services firms watched customers get locked out of accounts. Transactions failed. Loan origination queues froze.
It was, in effect, an unannounced business continuity drill — and the results showed who had actually thought through their cloud dependency and who had written “cloud provider SLAs” in their BCP and called it a day.
TL;DR
- The October 20, 2025 AWS outage was a 15-hour disruption affecting 1,000+ companies, including financial institutions that couldn’t process transactions or serve customers
- The OCC’s 2026 Cybersecurity and Financial System Resilience Report names cloud concentration risk as a systemic concern; the Spring 2026 Semiannual Risk Perspective flags it as a current operational risk priority
- The FFIEC BCM Handbook requires institutions to identify single-provider dependencies and document alternative arrangements for critical outsourced technology services
- Most fintech BCPs understate cloud risk by treating provider SLAs as a resilience substitute rather than mapping actual RTOs to a provider-down scenario
- Fourth-party cloud concentration — when your vendors all run on the same hyperscaler — is a gap almost nobody maps
What Actually Happened on October 20
The AWS outage started in AWS’s us-east-1 region (northern Virginia) — the largest AWS region globally and home to a disproportionate share of financial services workloads. A malfunction in an internal subsystem that monitors the health of network load balancers triggered DNS resolution failures that spread well beyond the initial fault point.
The financial impact was immediate: financial platforms couldn’t authenticate users, payment processing queues backed up, and customer-facing applications that depended on AWS-hosted APIs went dark. Some institutions experienced partial recovery within hours; others were disrupted for the full 15-hour duration.
The scale matters. Over 6.5 million user reports during the incident. More than a thousand companies across industries, but financial services firms disproportionately affected because of the concentration of fintech and banking infrastructure in us-east-1. The OCC’s 2026 Cybersecurity and Financial System Resilience Report subsequently identified this as a concrete illustration of what technology concentration risk looks like when it materializes — not a theoretical scenario but an event that actually disrupted financial services at scale.
The Regulatory Picture Before and After October 2025
Regulators didn’t need October 2025 to tell them cloud concentration risk was a problem. The risk had been named, documented, and flagged in supervisory guidance for years. What October 2025 did was convert an abstract risk into a recent, memorable enforcement data point.
The FFIEC Business Continuity Management Handbook — the primary BCM examination framework — addresses technology service provider concentration explicitly. The guidance notes that when multiple critical functions depend on a single TSP, that dependency creates a potential single point of failure. Financial institutions are required to assess alternative providers or in-house arrangements for critical outsourced services, and to verify that critical TSPs have the capacity and resilience to meet the institution’s BCM needs.
The OCC’s Spring 2026 Semiannual Risk Perspective named operational resilience — including technology concentration — as a key current risk area. The context is clear: examiners reviewing BCP programs in 2026 are not reading cloud concentration risk as a hypothetical. They’re reading it in the context of an event that happened eight months ago.
Treasury established its Cloud Executive Steering Group in 2023 to coordinate across agencies on financial sector cloud concentration. That work has continued — and the October 2025 outage accelerated the supervisory focus on what financial institutions have actually done in response.
What the FFIEC Standard Actually Requires
The FFIEC BCM guidance creates several specific obligations for technology concentration risk that most fintech BCPs don’t fully address:
Dependency Mapping That Reaches Cloud Providers
Your BCP is supposed to identify critical business processes and map their dependencies — systems, vendors, data, staff. The cloud provider is a dependency. If your BIA says your loan origination system depends on your origination platform vendor but doesn’t trace that vendor’s dependency on AWS us-east-1, you’ve mapped one layer and stopped.
FFIEC examiners reviewing BCPs in 2026 are asking whether the dependency map reaches the cloud hyperscaler level. Not “what cloud vendor does your processor use” as an academic question — but whether that cloud dependency is accounted for in your RTO calculations and your recovery scenarios.
RTOs That Reflect Provider Dependency
If your critical system has an RTO of 4 hours but it runs on a single cloud provider that has just demonstrated a 15-hour outage capacity, your stated RTO is unrealistic unless you have a tested cross-region or cross-provider failover capability.
The FFIEC guidance requires that RTOs be achievable. A 4-hour RTO for a system with no tested failover path from a provider disruption is a documentation problem waiting to become an exam finding. The question examiners ask is: how do you achieve this RTO if your primary cloud environment is down for longer than the RTO period?
Alternative Arrangements for Critical Services
The guidance requires that where an alternative TSP is not readily available, the institution should evaluate options to continue business operations. For cloud-dependent services, this translates into three possible approaches:
- Multi-region deployment: Critical workloads are deployed in at least two cloud regions, with tested automatic or rapid-manual failover
- Multi-cloud deployment: Critical workloads have a parallel capability with a second provider — architecturally complex and expensive, but the most resilient option
- Documented tolerance: If neither approach is in place, the BCP explicitly documents that the institution accepts the risk of extended provider outage and has communicated that tolerance to its board and bank partner
Option 3 is honest — but it requires documentation of what happens during an extended outage, how customers are communicated with, what regulatory notifications are required, and at what point the disruption crosses into a reportable computer-security incident.
Testing Against Realistic Scenarios
The BCM handbook requires that BCP testing include scenarios relevant to the institution’s actual risk profile. An institution whose critical infrastructure is concentrated on a single cloud provider needs to test a provider-down scenario — not just a “our servers are offline” scenario.
Testing in this context means a tabletop exercise that walks through the October 20 scenario: the provider is down for an unknown duration, you don’t know when it will recover, your customers can’t log in, your transactions are failing. What are the escalation steps? Who notifies the bank partner? What’s the customer communication? When does this become a reportable incident under the OCC and FDIC 36-hour computer-security incident notification requirement?
The Fourth-Party Problem Nobody Is Mapping
Here’s the gap that appears in almost every TPRM and BCP program when examined closely: the vendor list shows multiple vendors in each critical function area, which looks like redundancy. But all three vendors run on AWS.
That’s not redundancy. That’s four contracts pointing at the same single point of failure.
As covered in our analysis of TPRM exam findings in 2026, fourth-party risk — your vendors’ infrastructure dependencies — is increasingly on the FFIEC examination agenda. The TPRM interagency guidance specifically requires institutions to understand their critical third parties’ subcontracting practices, and cloud hosting is one of the most commonly missed subcontracting dependencies.
The question to ask for every critical vendor in your inventory: what cloud provider hosts their production environment? If the answer is the same for three or more critical vendors, you have a concentration risk that your BCP needs to address regardless of how many separate vendor contracts you have.
This was part of why the October 2025 outage hit financial services harder than the aggregate statistics suggested — many fintechs had built vendor redundancy that didn’t actually provide infrastructure redundancy when the underlying cloud provider went offline.
What Bank Partners Are Now Asking
Bank partners managing their own TPRM obligations are starting to add cloud concentration questions to fintech due diligence questionnaires. Specifically:
| Question | What They’re Actually Asking |
|---|---|
| What cloud providers do you use for production? | Are you a single-cloud shop? |
| What is your RTO for your primary payment processing flow? | Can you survive 24 hours of provider downtime? |
| Do you have a tested multi-region or cross-provider failover? | Have you tested this, or is it theoretical? |
| What are your critical vendors’ cloud hosting arrangements? | Have you mapped your fourth-party concentration? |
| How would you notify us if a cloud provider outage affected your service? | Who calls us, and when? |
The October 2025 event gave bank risk teams a concrete scenario to ask about — and a reasonable expectation that fintech partners have addressed it in their BCPs since then.
The BCP Update Checklist
If you haven’t updated your BCP documentation to reflect cloud concentration risk since October 2025, here’s the gap analysis to run:
BIA layer: Does your BIA identify which critical processes run on which cloud providers — not just which vendors, but which providers power those vendors?
RTO achievability: For every critical process with an RTO under 24 hours, do you have a tested recovery path that works when your primary cloud environment is offline? If not, is the RTO realistic?
Scenario coverage: Does your tabletop exercise library include a “primary cloud provider down for unknown duration” scenario? If not, add it.
Fourth-party mapping: For your top five to ten critical vendors, have you confirmed their cloud hosting arrangements and assessed whether your vendor diversification provides actual infrastructure diversification?
Communication runbooks: Do your bank partner notification templates and customer communication templates include the cloud provider outage scenario specifically?
Disclosure threshold: Have you documented when a cloud-related service disruption crosses into the reportable incident category under your federal banking regulator notification obligations?
The DORA operational resilience requirements — covered in our analysis of the first ICT incident report — add a parallel European lens on concentration risk that’s relevant for any fintech serving EU customers.
So What?
The October 2025 AWS outage was not the first major cloud provider outage to hit financial services, and it won’t be the last. What changed is that it happened in the middle of an intensified OCC and FFIEC supervisory focus on technology concentration risk — and it gave regulators and bank partners a recent, concrete reference point for conversations about what your BCP actually covers.
BCPs that rely on provider SLAs as a resilience substitute, that state an RTO without a tested recovery path from provider downtime, or that map vendor dependencies without tracing those vendors to their cloud providers are increasingly going to look incomplete under examination.
The fix isn’t necessarily a cloud architecture overhaul. It’s an honest assessment of what your BCP says, what scenarios it’s been tested against, and whether your stated RTOs are achievable in a realistic provider-down scenario. In most cases, the gap is documentation and testing, not infrastructure.
That’s a week’s work. The Business Continuity & Disaster Recovery (BCP/DR) Kit includes a BIA template, worked examples for fintech and banking environments, tabletop exercise scenarios, and case study walkthroughs — including an AWS outage scenario — that can get your program to a defensible state.
For the TPRM side of the house — vendor dependency mapping, fourth-party risk assessment, and what your vendor contracts need to say about business continuity — see our TPRM exam findings analysis for what’s getting flagged today.
Sources: OCC 2026 Cybersecurity and Financial System Resilience Report | FFIEC BCM Handbook | OCC Spring 2026 Semiannual Risk Perspective | AWS October 2025 Outage Analysis — Financial Executives International
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What does the FFIEC Business Continuity Management handbook say about cloud concentration risk?
Does the FFIEC BCP guidance apply to fintechs that aren't directly chartered?
What is the OCC's 2026 assessment of cloud concentration risk in the financial sector?
What should a fintech's BCP say about cloud provider dependency?
What are examiners asking about cloud resilience in 2026 BCP reviews?
What is 'fourth-party cloud concentration' and why does it matter for fintech BCP?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Keep reading
Related posts.
Business Continuity
FFIEC BCM Section III.B Risk Assessment: Turn Threats Into Continuity Strategies
Build an FFIEC BCM Section III.B risk assessment that traces threats, controls, gaps, continuity strategies, tests, and remediation.
Jul 24, 2026
Business Continuity
BCP Testing That Actually Satisfies Examiners: What FFIEC Requires Beyond Your Annual Tabletop
An annual tabletop that never fails anything is not a BCP test — it's theater. Here's what the FFIEC Business Continuity Management booklet actually requires, how 2026 examiners evaluate test programs, and what documentation makes your tests defensible.
Jul 20, 2026
Business Continuity
Software Supply Chain Failure BCP Scenarios: What FFIEC Guidance and Post-CrowdStrike Expectations Require Financial Institutions to Test in 2026
A year after CrowdStrike brought down 8.5 million Windows devices and cost the banking sector $1.149 billion, financial institution BCP programs are expected to test software supply chain failure scenarios explicitly. Here's what changed in FFIEC guidance, what examiners now look for, and how to build a scenario that holds up.
Jul 12, 2026