Feature Business Continuity
AWS Went Down in October. Most BCPs Assumed It Wouldn't. Here's How to Fix That.
The October 2025 AWS DNS outage knocked out DynamoDB endpoints across multiple regions. Most financial institution BCPs treat cloud infrastructure as a given, not a dependency to plan around. Here's what FFIEC and DORA actually require — and what cloud-aware recovery planning looks like.
Table of Contents
TL;DR
- AWS experienced a major DNS resolution outage in October 2025 affecting DynamoDB endpoints across regions; Azure had a Front Door global failure that took down Microsoft 365 and affected enterprise customers including Alaska Airlines. Both exposed the same gap: most financial institution BCPs don’t plan for cloud provider failure — they plan for internal systems failing while assuming cloud is always up.
- The FFIEC Business Continuity Management booklet requires your BIA to explicitly model cloud dependencies with recovery time and recovery point objectives for cloud-unavailable scenarios. Assuming AWS or Azure is always available is an examination finding.
- OCC Bulletin 2023-17 treats cloud providers as critical third parties requiring full TPRM documentation, contract review, and ongoing monitoring — not just a vendor agreement with an SLA.
- DORA (effective January 2025 for EU financial entities) adds concentration risk assessment requirements for hyperscaler dependencies. US fintechs with EU operations are already on the hook.
When the October 2025 AWS outage hit, most financial institutions discovered two things simultaneously: their core systems were down, and their business continuity plans had nothing useful to say about it.
The outage traced to a DNS resolution failure affecting DynamoDB database service endpoints across multiple AWS regions. Payment processing queues stalled. Customer-facing apps threw errors. Backend workflows that depended on DynamoDB for state management froze. It wasn’t an infrastructure failure in the traditional sense — physical servers were fine. The network path that told your application where to find its database stopped working.
Institutions that responded well had done something unusual: they had modeled cloud provider outages in their Business Impact Analysis. They had documented degraded operating procedures for cloud-unavailable states. They had tested — actually tested, not tabletop-discussed — what it looked like to operate for four hours without their primary cloud provider.
Institutions that struggled had BCPs that treated “the cloud” as infrastructure, not a dependency with its own failure modes. Their recovery procedures started from the assumption that cloud services were available. Until October 2025, that assumption had always held.
Why BCPs Get Cloud Wrong
Cloud changed the risk profile of operational resilience without changing how most organizations think about it. The traditional BCP mental model: your on-premises data center goes down, your DR site in a different geography comes online, your staff switches over, you’re running again. Cloud was supposed to make this easier — no DR site needed, everything redundant, always-on infrastructure.
That model works for some failure modes. Not for the ones that actually happened in 2025.
The major hyperscalers each suffered at least one significant global-scale outage in 2025. AWS’s DNS failure in October. An Azure Front Door failure that knocked out Microsoft 365 — Teams, Outlook, enterprise authentication — affecting major enterprise customers including Alaska Airlines. A Google Cloud outage in the us-central1 region. These were failures in the cloud infrastructure itself, not local datacenter events your DR site can compensate for.
Most BCP templates ask: “What are your critical systems, and what’s the RTO if those systems fail?” The implicit assumption is that the compute and storage layer is available — you’re planning for application failures, database corruption, or edge network issues. Very few BCPs ask: “What’s your RTO if AWS us-east-1 is unavailable for 12 hours?”
That second question is now the one examiners are asking.
What the FFIEC Actually Requires
The FFIEC Business Continuity Management booklet is explicit. A Business Impact Analysis must identify and prioritize critical business functions and map their dependencies — including third-party dependencies like cloud infrastructure and SaaS platforms. The booklet doesn’t give cloud providers a pass because “that’s infrastructure.” Cloud dependencies are third-party dependencies, subject to the same BIA scrutiny as any other single point of failure.
This means your BIA needs answers to specific questions: Which critical functions depend on AWS, Azure, or GCP? Which region or availability zone? What’s the maximum tolerable downtime for each function if the cloud provider is down — not just if your application is down? What does “degraded operation” look like if cloud services are unavailable? Have you validated those degraded procedures?
If your BIA says “Core Banking: RTO 4 hours” but doesn’t specify whether that RTO assumes your cloud provider is available, your BIA has a gap. As covered in FFIEC BCM examination findings for 2026, examiners are specifically probing whether BIAs account for cloud and SaaS dependencies — it’s consistently among the top examination findings.
The 2020 FFIEC Joint Statement on Risk Management for Cloud Computing Services reinforced this: institutions using cloud services are responsible for managing the risks those services introduce, including availability risk. Cloud providers manage the underlying infrastructure; institutions manage the risk that infrastructure introduces to their operations.
Cloud Providers Are Critical Vendors Under OCC 2023-17
The interagency third-party risk management guidance — OCC Bulletin 2023-17, Federal Reserve SR 23-4, FDIC FIL 29-2023 — is unambiguous: technology service providers, including cloud providers, are third parties subject to your TPRM program when they support significant or critical activities.
That means your AWS or Azure agreement isn’t just a vendor contract. It’s a critical third-party relationship that should go through the same risk assessment, contract review, financial condition monitoring, and independent review that your core banking system vendor does.
In practice, this requires four things that most fintechs are missing:
Risk assessment: What would a cloud provider outage mean for each critical function? How long before impact is material to customers, regulators, or business operations? What’s your concentration — does your fintech’s entire stack run on one provider?
Contract review: Does your cloud agreement include SLAs with meaningful uptime commitments? Service credits if availability falls below threshold? Termination rights? Incident notification obligations that meet your regulatory response timelines?
Ongoing monitoring: Are you reviewing your cloud provider’s published service health dashboards on a defined cadence? Do you receive direct outage notifications? Is your monitoring sufficient for a critical third party or just a commodity vendor?
Independent review: Has someone outside your IT organization reviewed the adequacy of your cloud risk assessment and contract terms?
Most fintechs have AWS or Azure agreements signed by engineering. Most have never run those agreements through formal TPRM. That gap is exactly what OCC 2023-17 TPRM documentation reviews are surfacing in examinations.
What DORA Adds for EU-Connected Financial Institutions
If your fintech has EU operations, EU customers, or EU-based critical third parties, DORA — the Digital Operational Resilience Act, effective January 17, 2025 — applies to your ICT risk management framework.
DORA’s requirements go beyond traditional BCP in two important ways.
First, it mandates a comprehensive ICT third-party risk register that maps every dependency on cloud providers, payment processors, and critical software vendors. The register isn’t a TPRM tracker — it’s a structured mapping of dependencies with criticality classification, concentration risk analysis, and contingency documentation. This goes deeper than most US TPRM frameworks currently require. US fintechs with EU operations need to understand their DORA ICT register obligations and how they interact with FFIEC TPRM requirements.
Second, DORA explicitly addresses concentration risk at a systemic level: the risk that arises when the vast majority of the EU financial system depends on a small number of hyperscale cloud providers. European Supervisory Authorities now have direct oversight powers over designated critical ICT third-party providers — regulators can reach through your vendor relationship to assess the cloud provider itself.
For FFIEC’s own 2026 third-party guidance updates, concentration risk from cloud adoption is increasingly a named concern, with examiners evaluating whether institutions have assessed their exposure to hyperscaler concentration at both the firm level and the systemically interconnected level.
What a Cloud-Aware BCP Actually Requires
Six components make a BCP genuinely cloud-aware:
1. Cloud dependency mapping in the BIA. For each critical business function, document which cloud provider, which region or availability zone, and which specific services (compute, storage, database, networking) the function depends on. This is usually a worksheet expansion of your existing BIA, not a separate document.
2. Cloud-specific RTO/RPO targets. Set targets that account for cloud provider outages, not just application failures. If your payment processing RTO is “2 hours,” document whether that assumes your cloud provider is available. If it does and you can’t meet it without the cloud provider, you need a degraded-mode procedure or a revised RTO.
3. Degraded operating mode documentation. For each critical function that depends on cloud infrastructure, document what degraded operation looks like during a cloud outage. Can you process payments manually? Can you service existing customer inquiries? Can you submit required regulatory reports? These procedures need to be tested, not just written.
4. Multi-cloud or alternative path validation. If you’re “multi-cloud” in your architecture, verify that the redundancy is real. Many fintechs discovered during 2025 outages that their multi-cloud setup still routed DNS through a single provider, used a single CDN layer, or had database replication that depended on the failing region. Real redundancy means the backup path works independently.
5. Cloud provider TPRM documentation. Run your cloud agreements through your TPRM framework as you would any critical vendor — risk assessment, contract review, monitoring cadence, independent review, documented in your vendor management system.
6. Cloud outage tabletop exercise. Add a cloud provider outage scenario to your annual BCM testing calendar. Scenario: your primary cloud provider has been unavailable for 4 hours with no estimated restoration time. Walk through operational decisions, customer communications, regulatory reporting obligations, and escalation. The first time you answer these questions should not be during the actual outage. As cloud concentration risk analysis has found, many financial firms discover real gaps only when they try to run this scenario.
What Examiners Are Finding in 2026
Financial institution examiners reviewing BCMs in 2026 are consistently finding three cloud-related gaps:
BIAs that assume cloud availability. The BIA lists cloud services as dependencies but doesn’t model cloud-unavailable scenarios. RTOs are set without reference to cloud provider status.
Testing that never involves cloud failure. Annual tabletop exercises walk through hurricanes, power failures, and cyberattacks but never test cloud provider unavailability. Testing cloud failover means actually routing traffic away from the primary provider during a drill, not assuming the switch would work.
Missing TPRM for cloud agreements. Cloud providers are treated as utilities rather than critical vendors. No formal risk assessment, no compliance-lens contract review, no monitoring cadence documentation.
So What? The Practitioner Action List
If your BCP was written before 2024, it almost certainly has cloud-related gaps. Start here:
- Pull your BIA. Identify every critical function that lists a cloud provider or SaaS platform as a dependency. Count how many have cloud-conditional RTO targets versus cloud-assumed RTO targets.
- Pull your cloud agreements. Check for SLAs with financial remedies, incident notification obligations, and exit rights. Note what’s missing.
- Run your cloud providers through your TPRM framework — formally, with documentation. Treat AWS or Azure the same way you’d treat a core banking system vendor.
- Add a cloud outage scenario to your next BCM tabletop. The scenario doesn’t need to be elaborate: “AWS has been unavailable for 4 hours. Walk us through it.”
- Document degraded operating procedures for your top three cloud-dependent critical functions.
The Business Continuity & Disaster Recovery Kit includes a BIA template with explicit third-party and cloud dependency mapping, cloud-outage scenario guidance, and vendor risk documentation worksheets built to satisfy both FFIEC BCM and OCC 2023-17 requirements.
The October 2025 outage made a gap visible that was already there. Institutions that treat it as a wake-up call will close it before the next one. The ones that write it off are betting the cloud won’t fail again — which is a planning assumption, not a resilience strategy.
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
Does FFIEC guidance specifically address cloud provider outages in business continuity planning?
Is a cloud provider a 'critical vendor' under OCC Bulletin 2023-17?
What does DORA require for cloud concentration risk?
What should a cloud-aware Business Impact Analysis include?
What do examiners look for when reviewing cloud-related BCPs?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Keep reading
Related posts.
Business Continuity
Your BCP Is a Document. The FFIEC BCM Booklet Wants a Management Process. Here's What Examiners Are Testing.
The FFIEC Business Continuity Management booklet shifted the examination standard from recovery planning to operational resilience — but most fintechs and community banks still have a document, not a management process. Here are the seven BCM components, the most common examination findings, and what a defensible program actually looks like.
Sep 7, 2026
Business Continuity
DORA's ICT Register: Only 40% Filed Before the March Deadline. What Enforcement Looks Like Now.
DORA's ICT third-party register deadline passed on March 31, 2026. Only 40% of required entities submitted on time, and just 6.5% passed all quality checks. Here's what enforcement looks like — and why US fintechs that serve EU clients or provide cloud services can't treat this as someone else's problem.
Sep 3, 2026
Business Continuity
The Integration Requirement Your BCP Is Missing: What FFIEC Examiners Actually Check on Vendor Business Continuity
Collecting your vendor's SOC 2 and test summary isn't FFIEC BCM compliance. Examiners want to see that you've integrated your critical vendors' continuity plans into your own BCP—with evidence of end-to-end testing and notification tracking.
Aug 24, 2026