Skip to content
RiskTemplates · The Daily Brief Saturday, July 25, 2026
Wire FinCEN's Student Aid Fraud Alert: The ACH Refund Pattern Banks Need to Tune Now JUL 23

Feature Third-Party Risk

AI Vendor KRIs: Monitoring Model Drift, Output Errors, Complaints, and Contract Gaps

Six KRIs for monitoring AI vendor performance in your environment — model update detection, output error rate, complaint attribution, drift indicators, contract gap rate, and incident notification lag. Built around FS AI RMF and NIST AI RMF monitoring requirements.

By Rebecca Leung · May 30, 2026 ·
Table of Contents

TL;DR

  • Your fraud detection vendor silently updated their model in Q3. You found out when complaints about declined transactions spiked. The vendor called it an improvement. Your examiner will call it a third-party oversight gap
  • The Treasury’s Financial Services AI Risk Management Framework (FS AI RMF, March 2026) requires ongoing output-level monitoring of vendor AI performance — not annual questionnaire reviews and not point-in-time assessments
  • Six KRIs that catch what questionnaires miss: model update detection, output error rate in your environment, complaint attribution, drift indicators, contract gap rate, and vendor incident notification lag
  • “The vendor built it” has never been a regulatory defense — you remain responsible for what their AI does to your customers

Your fraud detection vendor silently updated their model in Q3. You found out when complaint volume about incorrectly declined transactions doubled in two weeks. The vendor’s explanation was that they’d improved accuracy across the full customer base. Your specific customer segment was collateral damage.

This scenario plays out constantly with AI vendors — not because vendors are dishonest, but because AI models in production are living systems. They get retrained, tuned, fine-tuned, replaced, and “improved” on schedules that have nothing to do with your compliance calendar. The model you assessed in your initial due diligence questionnaire six months ago may not be the model making decisions in your environment today.

Most TPRM programs treat AI vendor oversight like traditional software vendor oversight: annual questionnaire, SOC 2 report review, periodic meeting. That approach worked when vendor risk was about uptime, data handling, and contractual compliance. It doesn’t work when the vendor is making decisions about your customers in real time, and those decisions change whenever the model gets updated.

The FS AI RMF, released by the U.S. Treasury in March 2026 with 230 control objectives mapped across the AI lifecycle, is explicit about this gap: financial institutions using third-party AI are required to implement ongoing monitoring — output-level performance tracking, change detection, and incident response capabilities — not annual snapshots. The framework treats vendor AI the same as internally developed AI for monitoring purposes, because the risk to your customers is identical regardless of who built the model.

Here are the six KRIs that close the monitoring gap.


Why Traditional Vendor Monitoring Fails for AI

Traditional TPRM monitoring asks: Is the vendor financially stable? Are their SOC 2 controls working? Are they complying with contractual data handling requirements?

These questions matter. They’re also insufficient for AI vendors, for a specific reason: the product you’re buying is a model whose outputs change over time. You’re not buying a static system that either works or doesn’t. You’re buying a decisioning system that produces different outputs as data patterns change, as the vendor retrains, and as the model drifts from the environment it was originally optimized for.

The GAO’s 2025 report on AI use and oversight in financial services identified third-party AI as one of the most significant governance gaps across regulated institutions — specifically citing the absence of ongoing performance monitoring for vendor AI systems as a consistent finding. Institutions could demonstrate due diligence at point of onboarding but not continuous oversight of model behavior in production.

The six KRIs below address the specific failure modes that questionnaires and SOC 2 reviews can’t detect.


The 6 AI Vendor KRIs

KRI 1: Unannounced Model Update Detection Rate

What it measures: The percentage of material model updates by your AI vendor that were detected through your monitoring program versus disclosed proactively by the vendor.

Why it matters: Most AI vendors update models continuously in production without formal notification. Unless your contract requires advance notice of material updates, you’ll find out about changes when outputs shift, complaints spike, or a bias testing routine surfaces a new pattern. This KRI measures your detection capability — how many updates are you catching, and how quickly?

Data source: Baseline output distribution established at deployment and tracked monthly. Model update indicators include: output volume shifts, latency changes, behavioral changes on identical test inputs (requires a canary test set), and vendor release communications.

ThresholdCriteria
Green>90% of material model updates detected within 30 days; vendor proactively notifies for >50% of material updates
AmberDetection rate 70–90%; or average detection lag >30 days; or no contractual notification requirement in place
RedDetection rate <70%; or detection lag >60 days for any material update; or no baseline output distribution maintained

Owner: AI governance or model risk function. If neither exists, TPRM with technical support from product or engineering to establish baseline distributions.


KRI 2: AI Output Error Rate (Measured in Your Environment)

What it measures: The rate at which the vendor’s AI produces outputs that your downstream processes flag as incorrect, inconsistent with expected behavior, or requiring manual review — measured in your environment, not in the vendor’s testing data.

Why it matters: Vendor-reported accuracy metrics reflect performance on the vendor’s test data in the vendor’s evaluation environment. Your customers, your data distribution, and your use case are different. The only error rate that matters for your compliance program is the one measured in your actual deployment.

The CFPB’s August 2022 action against Hello Digit — a $2.7M penalty plus consumer restitution — involved an AI system whose vendor-claimed accuracy coexisted with approximately 70,000 overdraft reimbursement requests from customers harmed by the algorithm’s actual behavior in production. The error rate in the vendor’s testing environment was not the error rate in the customer’s experience.

Data source: Post-decision review logs, manual override records, downstream exception flags, customer service escalations tied to AI-driven decisions.

ThresholdCriteria
GreenOutput error rate within ±20% of contractual performance threshold; no errors in protected-class or regulatory-sensitive decision categories
AmberError rate 20–40% above contractual threshold; or error pattern concentrated in a specific customer segment, product, or decision type
RedError rate >40% above contractual threshold; or any error pattern with potential fair lending, UDAAP, or adverse action implication

Owner: Product or operations team running the process that relies on the AI output, with second-line AI governance or compliance reviewing the error pattern analysis.


KRI 3: Customer Complaint Rate Attributed to AI Vendor Output

What it measures: The percentage of customer complaints that, upon root cause analysis, trace to the AI vendor’s model output as the proximate cause — incorrect decision, unexpected behavior, or failure to perform as marketed.

Why it matters: Customer complaints are the leading consumer-harm signal that regulators look at first. If complaints about declined transactions, credit decisions, or account actions are rising, and the common thread in root cause analysis is vendor AI behavior, you have a UDAAP-adjacent situation developing whether or not a formal error rate threshold is being breached.

The OCC’s model risk guidance — updated in 2026 to replace SR 11-7 — specifically requires financial institutions to maintain complaint monitoring for AI-driven decisions and to trace complaints to model-level root causes, not just front-line resolution.

Data source: Complaint management system with root cause attribution tagging. Requires a taxonomy that distinguishes AI-vendor-output complaints from pricing disputes, service complaints, and other categories.

ThresholdCriteria
GreenAI-attributed complaints constitute <3% of total complaint volume; no complaint cluster traceable to a specific model output or decision type
Amber3–8% of complaints attributed to AI vendor output; or any cluster of ≥5 complaints sharing a common AI root cause in a single period
Red>8% AI-attributed; or any complaint with fair lending, disparate impact, or adverse action implication traceable to vendor model output

Owner: Compliance or consumer protection function, with AI governance providing root cause analysis support.


KRI 4: Output Distribution Shift (Drift Proxy)

What it measures: The degree to which the vendor’s AI output distribution — approval rates, score distributions, flagging rates, classification frequencies — has shifted from the established baseline, indicating potential model drift or undisclosed model changes.

Why it matters: You can’t directly observe whether a vendor’s model has drifted. You can observe what the model does in your environment. A fraud detection model that suddenly flags 35% more transactions than it did last quarter hasn’t necessarily gotten better — it may have been retrained on different data, or the real-world data distribution has shifted enough to degrade performance on your customer base.

Output distribution shift is the most reliable indirect signal of model change or drift, and the one most consistently identified in the FS AI RMF’s monitoring requirements as a required ongoing control.

Data source: Monthly tracking of output distribution statistics — approval/denial rate by product, score distribution, flag rate by decision category — compared against the baseline established at deployment.

ThresholdCriteria
GreenOutput distribution within ±15% of baseline across all decision categories
AmberDistribution shift of 15–30% in any decision category; or shift sustained across two consecutive periods; or vendor unable to explain shift when queried
RedDistribution shift >30% in any decision category; or any shift correlated with a demographic or segment pattern that warrants fair lending review

Owner: AI governance or model risk, with data provided by the team running the AI system. Monthly monitoring cadence is the minimum; weekly for high-volume, high-stakes use cases.


KRI 5: Contract Gap Rate

What it measures: The percentage of active AI vendor contracts that are missing one or more of the material governance terms required to support effective monitoring — quantified performance thresholds, bias testing obligations, incident notification timelines, audit rights, model change notification requirements, and termination triggers.

Why it matters: You can’t monitor against a threshold that doesn’t exist, escalate based on an incident notification SLA that isn’t in the contract, or invoke audit rights you never negotiated. Contract gaps are governance gaps. And in AI vendor relationships, contract gaps are remarkably common — most vendor contracts were written for traditional software, not for a model that makes real-time decisions about your customers and gets updated without your input.

A 2026 analysis of FS AI RMF implementation gaps found that AI vendor contracts at financial institutions frequently lacked the key terms needed for ongoing oversight: performance benchmarks tied to your customer base rather than vendor-wide averages, contractual notification for material model updates, and explicit audit rights for third-party bias testing.

Data source: Contract management system with a vendor AI governance term checklist applied to all AI-relevant vendor contracts.

ThresholdCriteria
Green>90% of AI vendor contracts contain all required governance terms; any gaps have a documented remediation plan and target date
Amber80–90% complete; or any Critical/High-tier AI vendor contract missing incident notification or audit rights terms
Red<80% complete; or any AI vendor contract that’s been flagged as requiring renegotiation for >180 days without resolution

Owner: TPRM or legal, with AI governance providing the required-term checklist. Contract gap remediation should be part of every AI vendor contract renewal cycle.


KRI 6: Vendor Incident Notification Lag

What it measures: The elapsed time between a vendor AI incident occurring and the vendor’s formal notification to your organization — measured against contractual SLAs where they exist, and tracked absolutely where they don’t.

Why it matters: When your AI vendor experiences a model failure, a data issue, or a security event affecting the model, the timing of your notification directly affects your ability to contain customer harm, meet your own regulatory notification obligations, and take corrective action before the problem compounds.

The FFIEC’s 36-hour notification requirement for computer security incidents requires banking organizations to notify their primary federal regulator within 36 hours. If your AI vendor notifies you on day three that they had a model failure on day one, you may already be in violation of your own reporting obligation — for an event you didn’t cause and couldn’t have known about without timely vendor notification.

Data source: Vendor incident log with notification timestamps compared against incident date (where available from vendor) or symptom-detection date (where incident date is unknown).

ThresholdCriteria
Green100% of vendor AI incidents notified within contractual SLA; contractual SLA ≤24 hours for Critical-tier AI vendors
AmberAny incident notified outside contractual SLA; or no contractual notification SLA in place for any Critical/High-tier AI vendor
RedAny notification lag that caused your organization to miss or risk missing a regulatory notification obligation; or vendor failed to notify for a known incident

Owner: Vendor relationship owner, with escalation to compliance and legal any time a notification lag creates regulatory reporting exposure.


Summary Table

KRIPrimary SignalData SourceReview Cadence
Unannounced Model Update DetectionSilent model changesOutput baseline + vendor commsMonthly
AI Output Error RateIn-environment performance gapReview logs, override recordsMonthly
Complaint Attribution to AI OutputConsumer harm signalsComplaint system + root cause taggingMonthly
Output Distribution ShiftDrift and undisclosed model changesMonthly distribution statisticsMonthly
Contract Gap RateGovernance term completenessContract management systemQuarterly
Vendor Incident Notification LagDelayed regulatory exposureIncident log with notification timestampsPer incident

Connecting AI Vendor KRIs to Your Broader Risk Program

AI vendor KRIs should feed two other parts of your risk program. First, AI model drift and hallucination KRIs — if your AI vendor output distribution is shifting, the model-level drift KRI should be receiving the same data and triggering parallel escalation. The two programs should share data, not duplicate effort.

Second, your standard vendor due diligence review process should be updated to incorporate AI-specific monitoring evidence as part of the annual review. A vendor that has a clean SOC 2 and no financial stability concerns but a deteriorating output error rate and unannounced model updates is a high-risk vendor by any risk-based assessment, even if the traditional TPRM metrics are all green.

The critical vendor monitoring framework applies to AI vendors with equal force — and the stakes are higher because a critical AI vendor failure isn’t just an operational disruption. It’s a decisioning failure that may have already produced customer harm before your monitoring caught it.


The Accountability Problem AI Vendor Monitoring Solves

The CFPB’s position — stated in its August 2024 comment to the Treasury Department — is unambiguous: “There are no exceptions to the federal consumer financial protection laws for new technologies.” This applies to vendor AI just as directly as it applies to your own models.

If your fraud detection vendor’s model flags Spanish-speaking customers at twice the rate of English-speaking customers on identical transaction patterns, that’s your fair lending problem. If your credit decisioning vendor silently updates their model and your denial rates in lower-income zip codes double the following month, that’s your ECOA problem. If your customer service AI starts generating inaccurate fee disclosures after a model update, that’s your UDAAP problem.

The six KRIs above are what “we monitor vendor AI performance” looks like in practice — not what it looks like in a questionnaire response or a due diligence checklist.


So What?

AI vendor monitoring is the gap between point-in-time due diligence and continuous accountability. The questionnaire you sent at onboarding described the model as it existed when the relationship started. The KRIs above describe what the model is doing to your customers today.

If you can’t answer the following from existing data, that’s where to start:

  • Has your AI vendor’s output distribution shifted more than 15% in any decision category in the last 90 days?
  • How many customer complaints from last quarter traced to AI vendor output in root cause analysis?
  • Which of your Critical-tier AI vendor contracts is missing a contractual incident notification SLA?

The Third-Party Risk Management (TPRM) Kit includes an AI-specific vendor due diligence questionnaire, ongoing monitoring templates for Critical and High-tier AI vendors, and contract gap checklists aligned to FS AI RMF and OCC 2026 model risk guidance.

◆ Need the working template?

Start with the source guide.

These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.

◆ Immaterial Findings · Weekly

Sharp risk & compliance insights. No fluff.

◆ FAQ

Frequently asked questions.

What is an AI vendor KRI?
An AI vendor KRI is a metric that signals when a third-party AI model you've deployed is drifting in performance, producing errors, generating complaints, or operating outside the parameters your contract specifies — before those issues become customer harm, regulatory findings, or enforcement actions. Key AI vendor KRIs include: unannounced model update detection rate, output error rate measured in your environment, customer complaint rate attributed to AI decisions, output distribution shift as a drift proxy, contract gap rate (missing performance thresholds, audit rights, incident notification SLAs), and vendor incident notification lag.
Do I need to monitor AI vendor performance if I'm just a customer using their tool?
Yes — and regulators are explicit about it. The Treasury's Financial Services AI Risk Management Framework (FS AI RMF, released March 2026) specifies ongoing output-level monitoring as an obligation for financial institutions using third-party AI, not just annual questionnaire reviews. The CFPB's position is equally clear: there are no exceptions to consumer financial protection laws for new technologies. If the vendor's AI produces a harmful outcome for your customer, it's your compliance problem regardless of whether you built the model.
How do I detect when a vendor has silently updated their AI model?
Direct disclosure from the vendor is rare without contractual requirements — most AI vendors update models continuously in production without formal notification. Indirect signals include output distribution shifts (a fraud detection model that suddenly flags 40% more transactions than last month), latency changes, behavioral differences on identical inputs across time, and vendor release notes (if contractually required). Your monitoring program should establish baseline output distributions when you first deploy and track deviation from that baseline as the primary drift detection signal.
What should be in an AI vendor contract to support KRI monitoring?
At minimum: quantified performance thresholds (accuracy, error rates, false positive/negative rates); bias testing obligations by protected class; incident notification timelines for harmful outputs; audit rights including independent testing; model change notification requirements (any material model update, with defined lead time); data handling restrictions for AI training; and termination triggers tied to specific performance failures. Contracts missing these terms create contract gap KRI exposure — you can't monitor against a contractual threshold that doesn't exist.
What do regulators expect for AI vendor oversight in 2026?
The FS AI RMF (Treasury, March 2026, 230 control objectives) treats ongoing monitoring of vendor AI performance as a core requirement — not a best practice. The framework requires output-level monitoring, performance benchmarking against contractual thresholds, change detection when vendor models are updated, and incident response capabilities when vendor AI produces harmful outputs. The OCC's 2026 model risk management guidance extends model risk requirements to vendor-supplied AI models as well as internally developed ones. You remain responsible for what the vendor's model does in your environment.
What is the CFPB Hello Digit case and why does it matter for AI vendor monitoring?
In August 2022, the CFPB took action against Hello Digit (acquired by Oportun) for an AI-driven savings algorithm that repeatedly caused the exact harm it promised to prevent. The algorithm was designed to transfer funds to savings without causing overdrafts — yet caused thousands of overdraft events that the company then failed to reimburse as promised. The CFPB imposed a $2.7M penalty plus consumer restitution. The case established that algorithmic output failures with consumer harm are UDAAP violations regardless of whether a human reviewed each decision — which is exactly what AI vendor monitoring KRIs are designed to catch before the CFPB does.
Rebecca Leung

Author

Rebecca Leung

Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.

◆ Related framework

Third-Party Risk Management (TPRM) Kit

Complete vendor risk management lifecycle from initial due diligence to ongoing oversight.

Immaterial Findings · Newsletter

The brief, in your inbox.

Enforcement of the week, a framework breakdown, and the prompts that are actually worth running. Delivered to your inbox. Free.