Feature Third-Party Risk
AI Vendor KRIs: Monitoring Model Drift, Output Errors, Complaints, and Contract Gaps
Six KRIs for monitoring AI vendor performance in your environment — model update detection, output error rate, complaint attribution, drift indicators, contract gap rate, and incident notification lag. Built around FS AI RMF and NIST AI RMF monitoring requirements.
Table of Contents
TL;DR
- Your fraud detection vendor silently updated their model in Q3. You found out when complaints about declined transactions spiked. The vendor called it an improvement. Your examiner will call it a third-party oversight gap
- The Treasury’s Financial Services AI Risk Management Framework (FS AI RMF, March 2026) requires ongoing output-level monitoring of vendor AI performance — not annual questionnaire reviews and not point-in-time assessments
- Six KRIs that catch what questionnaires miss: model update detection, output error rate in your environment, complaint attribution, drift indicators, contract gap rate, and vendor incident notification lag
- “The vendor built it” has never been a regulatory defense — you remain responsible for what their AI does to your customers
Your fraud detection vendor silently updated their model in Q3. You found out when complaint volume about incorrectly declined transactions doubled in two weeks. The vendor’s explanation was that they’d improved accuracy across the full customer base. Your specific customer segment was collateral damage.
This scenario plays out constantly with AI vendors — not because vendors are dishonest, but because AI models in production are living systems. They get retrained, tuned, fine-tuned, replaced, and “improved” on schedules that have nothing to do with your compliance calendar. The model you assessed in your initial due diligence questionnaire six months ago may not be the model making decisions in your environment today.
Most TPRM programs treat AI vendor oversight like traditional software vendor oversight: annual questionnaire, SOC 2 report review, periodic meeting. That approach worked when vendor risk was about uptime, data handling, and contractual compliance. It doesn’t work when the vendor is making decisions about your customers in real time, and those decisions change whenever the model gets updated.
The FS AI RMF, released by the U.S. Treasury in March 2026 with 230 control objectives mapped across the AI lifecycle, is explicit about this gap: financial institutions using third-party AI are required to implement ongoing monitoring — output-level performance tracking, change detection, and incident response capabilities — not annual snapshots. The framework treats vendor AI the same as internally developed AI for monitoring purposes, because the risk to your customers is identical regardless of who built the model.
Here are the six KRIs that close the monitoring gap.
Why Traditional Vendor Monitoring Fails for AI
Traditional TPRM monitoring asks: Is the vendor financially stable? Are their SOC 2 controls working? Are they complying with contractual data handling requirements?
These questions matter. They’re also insufficient for AI vendors, for a specific reason: the product you’re buying is a model whose outputs change over time. You’re not buying a static system that either works or doesn’t. You’re buying a decisioning system that produces different outputs as data patterns change, as the vendor retrains, and as the model drifts from the environment it was originally optimized for.
The GAO’s 2025 report on AI use and oversight in financial services identified third-party AI as one of the most significant governance gaps across regulated institutions — specifically citing the absence of ongoing performance monitoring for vendor AI systems as a consistent finding. Institutions could demonstrate due diligence at point of onboarding but not continuous oversight of model behavior in production.
The six KRIs below address the specific failure modes that questionnaires and SOC 2 reviews can’t detect.
The 6 AI Vendor KRIs
KRI 1: Unannounced Model Update Detection Rate
What it measures: The percentage of material model updates by your AI vendor that were detected through your monitoring program versus disclosed proactively by the vendor.
Why it matters: Most AI vendors update models continuously in production without formal notification. Unless your contract requires advance notice of material updates, you’ll find out about changes when outputs shift, complaints spike, or a bias testing routine surfaces a new pattern. This KRI measures your detection capability — how many updates are you catching, and how quickly?
Data source: Baseline output distribution established at deployment and tracked monthly. Model update indicators include: output volume shifts, latency changes, behavioral changes on identical test inputs (requires a canary test set), and vendor release communications.
| Threshold | Criteria |
|---|---|
| Green | >90% of material model updates detected within 30 days; vendor proactively notifies for >50% of material updates |
| Amber | Detection rate 70–90%; or average detection lag >30 days; or no contractual notification requirement in place |
| Red | Detection rate <70%; or detection lag >60 days for any material update; or no baseline output distribution maintained |
Owner: AI governance or model risk function. If neither exists, TPRM with technical support from product or engineering to establish baseline distributions.
KRI 2: AI Output Error Rate (Measured in Your Environment)
What it measures: The rate at which the vendor’s AI produces outputs that your downstream processes flag as incorrect, inconsistent with expected behavior, or requiring manual review — measured in your environment, not in the vendor’s testing data.
Why it matters: Vendor-reported accuracy metrics reflect performance on the vendor’s test data in the vendor’s evaluation environment. Your customers, your data distribution, and your use case are different. The only error rate that matters for your compliance program is the one measured in your actual deployment.
The CFPB’s August 2022 action against Hello Digit — a $2.7M penalty plus consumer restitution — involved an AI system whose vendor-claimed accuracy coexisted with approximately 70,000 overdraft reimbursement requests from customers harmed by the algorithm’s actual behavior in production. The error rate in the vendor’s testing environment was not the error rate in the customer’s experience.
Data source: Post-decision review logs, manual override records, downstream exception flags, customer service escalations tied to AI-driven decisions.
| Threshold | Criteria |
|---|---|
| Green | Output error rate within ±20% of contractual performance threshold; no errors in protected-class or regulatory-sensitive decision categories |
| Amber | Error rate 20–40% above contractual threshold; or error pattern concentrated in a specific customer segment, product, or decision type |
| Red | Error rate >40% above contractual threshold; or any error pattern with potential fair lending, UDAAP, or adverse action implication |
Owner: Product or operations team running the process that relies on the AI output, with second-line AI governance or compliance reviewing the error pattern analysis.
KRI 3: Customer Complaint Rate Attributed to AI Vendor Output
What it measures: The percentage of customer complaints that, upon root cause analysis, trace to the AI vendor’s model output as the proximate cause — incorrect decision, unexpected behavior, or failure to perform as marketed.
Why it matters: Customer complaints are the leading consumer-harm signal that regulators look at first. If complaints about declined transactions, credit decisions, or account actions are rising, and the common thread in root cause analysis is vendor AI behavior, you have a UDAAP-adjacent situation developing whether or not a formal error rate threshold is being breached.
The OCC’s model risk guidance — updated in 2026 to replace SR 11-7 — specifically requires financial institutions to maintain complaint monitoring for AI-driven decisions and to trace complaints to model-level root causes, not just front-line resolution.
Data source: Complaint management system with root cause attribution tagging. Requires a taxonomy that distinguishes AI-vendor-output complaints from pricing disputes, service complaints, and other categories.
| Threshold | Criteria |
|---|---|
| Green | AI-attributed complaints constitute <3% of total complaint volume; no complaint cluster traceable to a specific model output or decision type |
| Amber | 3–8% of complaints attributed to AI vendor output; or any cluster of ≥5 complaints sharing a common AI root cause in a single period |
| Red | >8% AI-attributed; or any complaint with fair lending, disparate impact, or adverse action implication traceable to vendor model output |
Owner: Compliance or consumer protection function, with AI governance providing root cause analysis support.
KRI 4: Output Distribution Shift (Drift Proxy)
What it measures: The degree to which the vendor’s AI output distribution — approval rates, score distributions, flagging rates, classification frequencies — has shifted from the established baseline, indicating potential model drift or undisclosed model changes.
Why it matters: You can’t directly observe whether a vendor’s model has drifted. You can observe what the model does in your environment. A fraud detection model that suddenly flags 35% more transactions than it did last quarter hasn’t necessarily gotten better — it may have been retrained on different data, or the real-world data distribution has shifted enough to degrade performance on your customer base.
Output distribution shift is the most reliable indirect signal of model change or drift, and the one most consistently identified in the FS AI RMF’s monitoring requirements as a required ongoing control.
Data source: Monthly tracking of output distribution statistics — approval/denial rate by product, score distribution, flag rate by decision category — compared against the baseline established at deployment.
| Threshold | Criteria |
|---|---|
| Green | Output distribution within ±15% of baseline across all decision categories |
| Amber | Distribution shift of 15–30% in any decision category; or shift sustained across two consecutive periods; or vendor unable to explain shift when queried |
| Red | Distribution shift >30% in any decision category; or any shift correlated with a demographic or segment pattern that warrants fair lending review |
Owner: AI governance or model risk, with data provided by the team running the AI system. Monthly monitoring cadence is the minimum; weekly for high-volume, high-stakes use cases.
KRI 5: Contract Gap Rate
What it measures: The percentage of active AI vendor contracts that are missing one or more of the material governance terms required to support effective monitoring — quantified performance thresholds, bias testing obligations, incident notification timelines, audit rights, model change notification requirements, and termination triggers.
Why it matters: You can’t monitor against a threshold that doesn’t exist, escalate based on an incident notification SLA that isn’t in the contract, or invoke audit rights you never negotiated. Contract gaps are governance gaps. And in AI vendor relationships, contract gaps are remarkably common — most vendor contracts were written for traditional software, not for a model that makes real-time decisions about your customers and gets updated without your input.
A 2026 analysis of FS AI RMF implementation gaps found that AI vendor contracts at financial institutions frequently lacked the key terms needed for ongoing oversight: performance benchmarks tied to your customer base rather than vendor-wide averages, contractual notification for material model updates, and explicit audit rights for third-party bias testing.
Data source: Contract management system with a vendor AI governance term checklist applied to all AI-relevant vendor contracts.
| Threshold | Criteria |
|---|---|
| Green | >90% of AI vendor contracts contain all required governance terms; any gaps have a documented remediation plan and target date |
| Amber | 80–90% complete; or any Critical/High-tier AI vendor contract missing incident notification or audit rights terms |
| Red | <80% complete; or any AI vendor contract that’s been flagged as requiring renegotiation for >180 days without resolution |
Owner: TPRM or legal, with AI governance providing the required-term checklist. Contract gap remediation should be part of every AI vendor contract renewal cycle.
KRI 6: Vendor Incident Notification Lag
What it measures: The elapsed time between a vendor AI incident occurring and the vendor’s formal notification to your organization — measured against contractual SLAs where they exist, and tracked absolutely where they don’t.
Why it matters: When your AI vendor experiences a model failure, a data issue, or a security event affecting the model, the timing of your notification directly affects your ability to contain customer harm, meet your own regulatory notification obligations, and take corrective action before the problem compounds.
The FFIEC’s 36-hour notification requirement for computer security incidents requires banking organizations to notify their primary federal regulator within 36 hours. If your AI vendor notifies you on day three that they had a model failure on day one, you may already be in violation of your own reporting obligation — for an event you didn’t cause and couldn’t have known about without timely vendor notification.
Data source: Vendor incident log with notification timestamps compared against incident date (where available from vendor) or symptom-detection date (where incident date is unknown).
| Threshold | Criteria |
|---|---|
| Green | 100% of vendor AI incidents notified within contractual SLA; contractual SLA ≤24 hours for Critical-tier AI vendors |
| Amber | Any incident notified outside contractual SLA; or no contractual notification SLA in place for any Critical/High-tier AI vendor |
| Red | Any notification lag that caused your organization to miss or risk missing a regulatory notification obligation; or vendor failed to notify for a known incident |
Owner: Vendor relationship owner, with escalation to compliance and legal any time a notification lag creates regulatory reporting exposure.
Summary Table
| KRI | Primary Signal | Data Source | Review Cadence |
|---|---|---|---|
| Unannounced Model Update Detection | Silent model changes | Output baseline + vendor comms | Monthly |
| AI Output Error Rate | In-environment performance gap | Review logs, override records | Monthly |
| Complaint Attribution to AI Output | Consumer harm signals | Complaint system + root cause tagging | Monthly |
| Output Distribution Shift | Drift and undisclosed model changes | Monthly distribution statistics | Monthly |
| Contract Gap Rate | Governance term completeness | Contract management system | Quarterly |
| Vendor Incident Notification Lag | Delayed regulatory exposure | Incident log with notification timestamps | Per incident |
Connecting AI Vendor KRIs to Your Broader Risk Program
AI vendor KRIs should feed two other parts of your risk program. First, AI model drift and hallucination KRIs — if your AI vendor output distribution is shifting, the model-level drift KRI should be receiving the same data and triggering parallel escalation. The two programs should share data, not duplicate effort.
Second, your standard vendor due diligence review process should be updated to incorporate AI-specific monitoring evidence as part of the annual review. A vendor that has a clean SOC 2 and no financial stability concerns but a deteriorating output error rate and unannounced model updates is a high-risk vendor by any risk-based assessment, even if the traditional TPRM metrics are all green.
The critical vendor monitoring framework applies to AI vendors with equal force — and the stakes are higher because a critical AI vendor failure isn’t just an operational disruption. It’s a decisioning failure that may have already produced customer harm before your monitoring caught it.
The Accountability Problem AI Vendor Monitoring Solves
The CFPB’s position — stated in its August 2024 comment to the Treasury Department — is unambiguous: “There are no exceptions to the federal consumer financial protection laws for new technologies.” This applies to vendor AI just as directly as it applies to your own models.
If your fraud detection vendor’s model flags Spanish-speaking customers at twice the rate of English-speaking customers on identical transaction patterns, that’s your fair lending problem. If your credit decisioning vendor silently updates their model and your denial rates in lower-income zip codes double the following month, that’s your ECOA problem. If your customer service AI starts generating inaccurate fee disclosures after a model update, that’s your UDAAP problem.
The six KRIs above are what “we monitor vendor AI performance” looks like in practice — not what it looks like in a questionnaire response or a due diligence checklist.
So What?
AI vendor monitoring is the gap between point-in-time due diligence and continuous accountability. The questionnaire you sent at onboarding described the model as it existed when the relationship started. The KRIs above describe what the model is doing to your customers today.
If you can’t answer the following from existing data, that’s where to start:
- Has your AI vendor’s output distribution shifted more than 15% in any decision category in the last 90 days?
- How many customer complaints from last quarter traced to AI vendor output in root cause analysis?
- Which of your Critical-tier AI vendor contracts is missing a contractual incident notification SLA?
The Third-Party Risk Management (TPRM) Kit includes an AI-specific vendor due diligence questionnaire, ongoing monitoring templates for Critical and High-tier AI vendors, and contract gap checklists aligned to FS AI RMF and OCC 2026 model risk guidance.
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Third-Party Risk Management (TPRM) Kit
Complete vendor risk management lifecycle from initial due diligence to ongoing oversight.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What is an AI vendor KRI?
Do I need to monitor AI vendor performance if I'm just a customer using their tool?
How do I detect when a vendor has silently updated their AI model?
What should be in an AI vendor contract to support KRI monitoring?
What do regulators expect for AI vendor oversight in 2026?
What is the CFPB Hello Digit case and why does it matter for AI vendor monitoring?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Third-Party Risk Management (TPRM) Kit
Complete vendor risk management lifecycle from initial due diligence to ongoing oversight.
◆ Keep reading
Related posts.
Third-Party Risk
Third-Party Risk Management Lifecycle RACI: Fix the Handoffs Between Procurement, Security, Legal, Business Owners, and Risk
A TPRM lifecycle RACI that assigns clear ownership at each stage — planning, due diligence, contracting, onboarding, monitoring, and offboarding — so findings don't fall between functions.
Jul 24, 2026
Third-Party Risk
Vendor Due Diligence Without a SOC 2: What Evidence Can Actually Substitute
A vendor due diligence checklist for evaluating security evidence when a vendor has no SOC 2 report, with a risk-based substitution matrix.
Jul 23, 2026
Third-Party Risk
Three Vendors, One Existential Risk: What the OCC's Community Bank Core Provider RFI Actually Asked
The OCC published Bulletin 2025-39 asking community banks hard questions about their relationships with Fiserv, FIS, and Jack Henry. The questions reveal exactly what examiners are now checking — and what most TPRM programs haven't addressed.
Jul 23, 2026