Skip to content
RiskTemplates · The Daily Brief Tuesday, September 15, 2026
Wire SEC's $64 Million Croft & Frost Offering Fraud Case: The Warning Email Compliance Teams Cannot Ignore SEP 14

Feature Business Continuity

AWS Went Down in October. Most BCPs Assumed It Wouldn't. Here's How to Fix That.

The October 2025 AWS DNS outage knocked out DynamoDB endpoints across multiple regions. Most financial institution BCPs treat cloud infrastructure as a given, not a dependency to plan around. Here's what FFIEC and DORA actually require — and what cloud-aware recovery planning looks like.

By Rebecca Leung · September 12, 2026 ·
Table of Contents

TL;DR

  • AWS experienced a major DNS resolution outage in October 2025 affecting DynamoDB endpoints across regions; Azure had a Front Door global failure that took down Microsoft 365 and affected enterprise customers including Alaska Airlines. Both exposed the same gap: most financial institution BCPs don’t plan for cloud provider failure — they plan for internal systems failing while assuming cloud is always up.
  • The FFIEC Business Continuity Management booklet requires your BIA to explicitly model cloud dependencies with recovery time and recovery point objectives for cloud-unavailable scenarios. Assuming AWS or Azure is always available is an examination finding.
  • OCC Bulletin 2023-17 treats cloud providers as critical third parties requiring full TPRM documentation, contract review, and ongoing monitoring — not just a vendor agreement with an SLA.
  • DORA (effective January 2025 for EU financial entities) adds concentration risk assessment requirements for hyperscaler dependencies. US fintechs with EU operations are already on the hook.

When the October 2025 AWS outage hit, most financial institutions discovered two things simultaneously: their core systems were down, and their business continuity plans had nothing useful to say about it.

The outage traced to a DNS resolution failure affecting DynamoDB database service endpoints across multiple AWS regions. Payment processing queues stalled. Customer-facing apps threw errors. Backend workflows that depended on DynamoDB for state management froze. It wasn’t an infrastructure failure in the traditional sense — physical servers were fine. The network path that told your application where to find its database stopped working.

Institutions that responded well had done something unusual: they had modeled cloud provider outages in their Business Impact Analysis. They had documented degraded operating procedures for cloud-unavailable states. They had tested — actually tested, not tabletop-discussed — what it looked like to operate for four hours without their primary cloud provider.

Institutions that struggled had BCPs that treated “the cloud” as infrastructure, not a dependency with its own failure modes. Their recovery procedures started from the assumption that cloud services were available. Until October 2025, that assumption had always held.

Why BCPs Get Cloud Wrong

Cloud changed the risk profile of operational resilience without changing how most organizations think about it. The traditional BCP mental model: your on-premises data center goes down, your DR site in a different geography comes online, your staff switches over, you’re running again. Cloud was supposed to make this easier — no DR site needed, everything redundant, always-on infrastructure.

That model works for some failure modes. Not for the ones that actually happened in 2025.

The major hyperscalers each suffered at least one significant global-scale outage in 2025. AWS’s DNS failure in October. An Azure Front Door failure that knocked out Microsoft 365 — Teams, Outlook, enterprise authentication — affecting major enterprise customers including Alaska Airlines. A Google Cloud outage in the us-central1 region. These were failures in the cloud infrastructure itself, not local datacenter events your DR site can compensate for.

Most BCP templates ask: “What are your critical systems, and what’s the RTO if those systems fail?” The implicit assumption is that the compute and storage layer is available — you’re planning for application failures, database corruption, or edge network issues. Very few BCPs ask: “What’s your RTO if AWS us-east-1 is unavailable for 12 hours?”

That second question is now the one examiners are asking.

What the FFIEC Actually Requires

The FFIEC Business Continuity Management booklet is explicit. A Business Impact Analysis must identify and prioritize critical business functions and map their dependencies — including third-party dependencies like cloud infrastructure and SaaS platforms. The booklet doesn’t give cloud providers a pass because “that’s infrastructure.” Cloud dependencies are third-party dependencies, subject to the same BIA scrutiny as any other single point of failure.

This means your BIA needs answers to specific questions: Which critical functions depend on AWS, Azure, or GCP? Which region or availability zone? What’s the maximum tolerable downtime for each function if the cloud provider is down — not just if your application is down? What does “degraded operation” look like if cloud services are unavailable? Have you validated those degraded procedures?

If your BIA says “Core Banking: RTO 4 hours” but doesn’t specify whether that RTO assumes your cloud provider is available, your BIA has a gap. As covered in FFIEC BCM examination findings for 2026, examiners are specifically probing whether BIAs account for cloud and SaaS dependencies — it’s consistently among the top examination findings.

The 2020 FFIEC Joint Statement on Risk Management for Cloud Computing Services reinforced this: institutions using cloud services are responsible for managing the risks those services introduce, including availability risk. Cloud providers manage the underlying infrastructure; institutions manage the risk that infrastructure introduces to their operations.

Cloud Providers Are Critical Vendors Under OCC 2023-17

The interagency third-party risk management guidance — OCC Bulletin 2023-17, Federal Reserve SR 23-4, FDIC FIL 29-2023 — is unambiguous: technology service providers, including cloud providers, are third parties subject to your TPRM program when they support significant or critical activities.

That means your AWS or Azure agreement isn’t just a vendor contract. It’s a critical third-party relationship that should go through the same risk assessment, contract review, financial condition monitoring, and independent review that your core banking system vendor does.

In practice, this requires four things that most fintechs are missing:

Risk assessment: What would a cloud provider outage mean for each critical function? How long before impact is material to customers, regulators, or business operations? What’s your concentration — does your fintech’s entire stack run on one provider?

Contract review: Does your cloud agreement include SLAs with meaningful uptime commitments? Service credits if availability falls below threshold? Termination rights? Incident notification obligations that meet your regulatory response timelines?

Ongoing monitoring: Are you reviewing your cloud provider’s published service health dashboards on a defined cadence? Do you receive direct outage notifications? Is your monitoring sufficient for a critical third party or just a commodity vendor?

Independent review: Has someone outside your IT organization reviewed the adequacy of your cloud risk assessment and contract terms?

Most fintechs have AWS or Azure agreements signed by engineering. Most have never run those agreements through formal TPRM. That gap is exactly what OCC 2023-17 TPRM documentation reviews are surfacing in examinations.

What DORA Adds for EU-Connected Financial Institutions

If your fintech has EU operations, EU customers, or EU-based critical third parties, DORA — the Digital Operational Resilience Act, effective January 17, 2025 — applies to your ICT risk management framework.

DORA’s requirements go beyond traditional BCP in two important ways.

First, it mandates a comprehensive ICT third-party risk register that maps every dependency on cloud providers, payment processors, and critical software vendors. The register isn’t a TPRM tracker — it’s a structured mapping of dependencies with criticality classification, concentration risk analysis, and contingency documentation. This goes deeper than most US TPRM frameworks currently require. US fintechs with EU operations need to understand their DORA ICT register obligations and how they interact with FFIEC TPRM requirements.

Second, DORA explicitly addresses concentration risk at a systemic level: the risk that arises when the vast majority of the EU financial system depends on a small number of hyperscale cloud providers. European Supervisory Authorities now have direct oversight powers over designated critical ICT third-party providers — regulators can reach through your vendor relationship to assess the cloud provider itself.

For FFIEC’s own 2026 third-party guidance updates, concentration risk from cloud adoption is increasingly a named concern, with examiners evaluating whether institutions have assessed their exposure to hyperscaler concentration at both the firm level and the systemically interconnected level.

What a Cloud-Aware BCP Actually Requires

Six components make a BCP genuinely cloud-aware:

1. Cloud dependency mapping in the BIA. For each critical business function, document which cloud provider, which region or availability zone, and which specific services (compute, storage, database, networking) the function depends on. This is usually a worksheet expansion of your existing BIA, not a separate document.

2. Cloud-specific RTO/RPO targets. Set targets that account for cloud provider outages, not just application failures. If your payment processing RTO is “2 hours,” document whether that assumes your cloud provider is available. If it does and you can’t meet it without the cloud provider, you need a degraded-mode procedure or a revised RTO.

3. Degraded operating mode documentation. For each critical function that depends on cloud infrastructure, document what degraded operation looks like during a cloud outage. Can you process payments manually? Can you service existing customer inquiries? Can you submit required regulatory reports? These procedures need to be tested, not just written.

4. Multi-cloud or alternative path validation. If you’re “multi-cloud” in your architecture, verify that the redundancy is real. Many fintechs discovered during 2025 outages that their multi-cloud setup still routed DNS through a single provider, used a single CDN layer, or had database replication that depended on the failing region. Real redundancy means the backup path works independently.

5. Cloud provider TPRM documentation. Run your cloud agreements through your TPRM framework as you would any critical vendor — risk assessment, contract review, monitoring cadence, independent review, documented in your vendor management system.

6. Cloud outage tabletop exercise. Add a cloud provider outage scenario to your annual BCM testing calendar. Scenario: your primary cloud provider has been unavailable for 4 hours with no estimated restoration time. Walk through operational decisions, customer communications, regulatory reporting obligations, and escalation. The first time you answer these questions should not be during the actual outage. As cloud concentration risk analysis has found, many financial firms discover real gaps only when they try to run this scenario.

What Examiners Are Finding in 2026

Financial institution examiners reviewing BCMs in 2026 are consistently finding three cloud-related gaps:

BIAs that assume cloud availability. The BIA lists cloud services as dependencies but doesn’t model cloud-unavailable scenarios. RTOs are set without reference to cloud provider status.

Testing that never involves cloud failure. Annual tabletop exercises walk through hurricanes, power failures, and cyberattacks but never test cloud provider unavailability. Testing cloud failover means actually routing traffic away from the primary provider during a drill, not assuming the switch would work.

Missing TPRM for cloud agreements. Cloud providers are treated as utilities rather than critical vendors. No formal risk assessment, no compliance-lens contract review, no monitoring cadence documentation.

So What? The Practitioner Action List

If your BCP was written before 2024, it almost certainly has cloud-related gaps. Start here:

  1. Pull your BIA. Identify every critical function that lists a cloud provider or SaaS platform as a dependency. Count how many have cloud-conditional RTO targets versus cloud-assumed RTO targets.
  2. Pull your cloud agreements. Check for SLAs with financial remedies, incident notification obligations, and exit rights. Note what’s missing.
  3. Run your cloud providers through your TPRM framework — formally, with documentation. Treat AWS or Azure the same way you’d treat a core banking system vendor.
  4. Add a cloud outage scenario to your next BCM tabletop. The scenario doesn’t need to be elaborate: “AWS has been unavailable for 4 hours. Walk us through it.”
  5. Document degraded operating procedures for your top three cloud-dependent critical functions.

The Business Continuity & Disaster Recovery Kit includes a BIA template with explicit third-party and cloud dependency mapping, cloud-outage scenario guidance, and vendor risk documentation worksheets built to satisfy both FFIEC BCM and OCC 2023-17 requirements.

The October 2025 outage made a gap visible that was already there. Institutions that treat it as a wake-up call will close it before the next one. The ones that write it off are betting the cloud won’t fail again — which is a planning assumption, not a resilience strategy.

◆ Need the working template?

Start with the source guide.

These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.

◆ Immaterial Findings · Weekly

Sharp risk & compliance insights. No fluff.

◆ FAQ

Frequently asked questions.

Does FFIEC guidance specifically address cloud provider outages in business continuity planning?
Yes. The FFIEC Business Continuity Management booklet requires institutions to include cloud providers and SaaS dependencies in their Business Impact Analysis. The BIA must identify which critical functions depend on cloud infrastructure, establish RTO and RPO targets for cloud-dependent services, and document what happens when the cloud provider is unavailable — not just when internal systems fail. Assuming cloud availability without modeling the failure mode is an examination finding.
Is a cloud provider a 'critical vendor' under OCC Bulletin 2023-17?
Yes. OCC Bulletin 2023-17 — the interagency third-party risk management guidance from 2023 — treats cloud providers as critical third parties when they support significant or critical activities. This means they require the same TPRM rigor as any core provider: risk assessment, contract review with SLA and termination rights, financial condition monitoring, and independent review. Most fintechs have cloud agreements but not cloud TPRM documentation.
What does DORA require for cloud concentration risk?
Under DORA (effective January 17, 2025 for EU financial entities), institutions must maintain an ICT third-party risk register that maps all dependencies on cloud and critical ICT providers. DORA also requires concentration risk assessment — evaluating what happens if a single cloud provider fails across the financial system. European Supervisory Authorities have direct oversight authority over designated critical ICT third-party providers. US fintechs with EU operations or customers face these requirements now.
What should a cloud-aware Business Impact Analysis include?
A cloud-aware BIA identifies every critical business function that depends on cloud infrastructure, maps which cloud provider and region it depends on, documents RTO and RPO for each function assuming cloud unavailability (not just internal failure), describes the degraded operating mode if the cloud provider is down, and assesses cross-cloud concentration. Many fintechs discover that their 'multi-cloud' architecture still routes through a single DNS provider or region.
What do examiners look for when reviewing cloud-related BCPs?
FFIEC examiners assess whether the BIA accounts for cloud dependencies, whether RTOs have been validated through actual cloud failover tests (not just tabletops), whether the institution has reviewed its cloud provider's own disaster recovery capabilities, whether cloud agreements include SLAs with service credit provisions, and whether there's a documented plan for degraded cloud availability. Most institutions have tested internal failover but never tested what happens when their primary cloud region is unavailable.
Rebecca Leung

Author

Rebecca Leung

Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.

◆ Related framework

Business Continuity & Disaster Recovery (BCP/DR) Kit

BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.

Immaterial Findings · Newsletter

The brief, in your inbox.

Enforcement of the week, a framework breakdown, and the prompts that are actually worth running. Delivered to your inbox. Free.