Feature Operational Risk
Process Risk KRIs: Backlogs, Manual Workarounds, SLA Breaches, and Error Rates
Six key risk indicators that catch operational process failures before they become loss events: backlog aging, manual workaround rate, SLA breach rate, error rate by process, rework volume, and handoff failure rate.
Table of Contents
TL;DR
- Process risk KRIs catch operational breakdowns before they become loss events or examiner findings — most programs track activity (items processed, tickets closed) rather than health signals like backlog aging and rework rate
- The OCC’s Fall 2025 Semiannual Risk Perspective identified operational risk as elevated, flagging institutions that rely on legacy systems and manual processes that weren’t designed to scale with current transaction volumes
- Six metrics that predict process failure: backlog aging by process type, manual workaround rate, SLA breach rate, error rate, rework volume, and handoff failure rate
- None of these require a GRC platform — they require a data owner, a threshold that means something, and an escalation path that actually gets used
Your operations manager calls it “how we get things done.” An examiner calls it a control failure waiting to become a loss event.
The problem with operational process risk isn’t that teams don’t know their processes are under stress. They usually know. Queues are building up, the same handoff fails every Tuesday afternoon, the reconciliation spreadsheet has 40 unresolved items from three weeks ago. The problem is that none of it appears in the risk dashboard — because nobody built a KRI for it.
Process risk is the category most KRI programs skip. It’s not as intuitive as fraud loss rates or liquidity metrics. It doesn’t map neatly to a single regulatory framework. The data lives in ticketing systems, workflow tools, or — in the worst case — the institutional memory of one overworked ops lead who’s leaving in two months.
But process failure is one of the most common drivers of operational losses in financial services. The BIS Working Paper on Operational and Cyber Risks in the Financial Sector found that execution, delivery, and process management failures consistently rank among the top drivers of operational loss events across financial institutions globally — accounting for a substantial share of both loss frequency and severity.
Here’s how to build the metrics that catch it before it becomes a loss.
Why Process Risk Goes Unmeasured
Most KRI programs are assembled backward — starting from a regulatory framework or whatever the GRC platform came loaded with, then working toward data availability. Process risk metrics get deprioritized because the data is messier (manual queues, informal tracking), the ownership is less clear (ops? first line? second line?), and the outputs look operational rather than strategic.
The result: dashboards with cyber KRIs, liquidity metrics, and compliance training completion rates — and nothing that captures whether the loan operations team’s reconciliation backlog has grown 40% over the past quarter.
The OCC’s Fall 2025 Semiannual Risk Perspective identified elevated operational risk, with particular concern about institutions relying on legacy infrastructure and manual processes that were never designed to scale with current transaction volumes. Examiners noted that institutions using workarounds as a substitute for system investment are accumulating process risk that isn’t showing up in their formal risk reporting.
That’s exactly the gap process risk KRIs are designed to close.
The 6 Process Risk KRIs That Actually Signal Failure
KRI 1: Backlog Aging by Process Type
What it measures: The volume and age of items waiting to be processed, segmented by process type — reconciliation queues, account maintenance requests, exception queues, dispute queues, document reviews, or whatever processes your institution tracks.
Why it matters: A backlog is normal. An aging backlog is a risk signal. The distinction is in the trend: is the queue clearing, holding steady, or growing? And how old is the oldest unresolved item?
A reconciliation queue with the same 35-day-old item sitting in it month after month isn’t just an ops problem — it’s a control failure. The item isn’t being resolved, and whatever risk it represents is accumulating without management action.
Data source: Ticketing systems, workflow queues, exception trackers. If none of these exist, the spreadsheet the ops team uses to track open items is your starting point. Start wherever the data is — perfect data infrastructure is not a prerequisite for monitoring.
| Threshold | Criteria |
|---|---|
| Green | No items older than 30 days in critical process queues; total queue volume within 10% of prior-month baseline |
| Amber | Any item older than 30 days in a critical queue; queue volume growing >15% month-over-month without documented spike explanation |
| Red | Items older than 60 days in any critical queue; queue volume growing >30% with no documented resolution plan |
Owner: Operations leads own the queue data. Risk or compliance sets the aging threshold and escalation trigger.
Calibration note: Define “critical” for your environment. Loan operations, payments clearing, regulatory reporting queues, and account reconciliation typically qualify. Internal HR ticket queues generally don’t.
KRI 2: Manual Workaround Rate
What it measures: The percentage of transactions, decisions, or handoffs in a given process that require a manual step not included in the designed process flow — typically because a system can’t handle the case, an integration is broken, or volume overwhelmed the automated path.
Why it matters: One documented manual workaround for a known system gap, with a named owner and a resolution timeline, is risk management. Twenty percent of transaction volume flowing through undocumented workarounds is a control failure operating invisibly.
Manual workarounds create risk in three ways. First, they’re not consistently logged — the step happens but no system records it. Second, they’re inconsistent — different people handle the same exception differently. Third, they’re silently permanent — what starts as a temporary patch while IT fixes a system issue quietly becomes the process.
Data source: Process walkthroughs, ops manager surveys, or exception logs. This metric often requires manual identification on first pass. The act of measuring it is itself a control — you’ll discover workarounds that nobody knew existed at the program level.
| Threshold | Criteria |
|---|---|
| Green | Documented workarounds constitute <5% of process volume; each has an owner, justification, and resolution timeline |
| Amber | 5–15% of volume relies on undocumented workarounds; or any workaround without documented ownership |
| Red | >15% undocumented; or any workaround active for >180 days without a documented system fix or process redesign |
Owner: Operations lead for each process, with second-line review to confirm documentation and compensating controls exist for active workarounds.
KRI 3: SLA Breach Rate
What it measures: The percentage of transactions, requests, or deliverables that miss their defined service level agreement — whether internal SLAs (ops team to business line) or external SLAs (institution to customer or counterparty).
Why it matters: SLAs exist because timing matters. A payment not processed on time isn’t just a customer experience problem — it’s a potential Reg E violation, a contractual breach, or a regulatory reporting miss. A compliance review running 45 days past its deadline isn’t just late — it’s a control failure with audit trail implications.
Most institutions track external SLAs (what they owe vendors or customers) and contractual commitments. Far fewer track internal SLAs — the commitments ops teams make to each other — which is exactly where process risk accumulates first and longest before anyone outside the team notices.
Data source: SLA tracking systems, operations dashboards, contract management tools. If internal SLAs aren’t currently documented, define them before building this KRI — the definition exercise alone will surface process risk you didn’t know you had.
| Threshold | Criteria |
|---|---|
| Green | SLA breach rate <5% across all process types; no breach in customer-facing or regulatory reporting categories |
| Amber | 5–15% breach rate in any non-critical process; or any single breach in a regulatory reporting or customer-facing commitment |
| Red | >15% in any category; or repeat breaches in the same process across two consecutive reporting periods |
Owner: Operations lead tracks the rate. Risk or compliance owns the regulatory and contractual SLA monitoring, including breach investigation and remediation.
KRI 4: Error Rate by Process Type
What it measures: The percentage of outputs from a given process that contain errors — failed entries, incorrect codes, incomplete documentation, miscalculations, or misrouted items — before any correction step is applied.
Why it matters: Error rate measures quality at the process level. A consistent 2% error rate looks manageable until your transaction volume triples — at which point you’re generating three times as many daily errors with the same team. Volume changes what a rate means, and that’s exactly the kind of shift that makes a static green threshold dangerously misleading.
Error rate also matters by process type. A 1% error rate in reconciliation and a 1% error rate in regulatory filings are not equivalent risk signals. The second one has a direct examination implication.
Data source: Quality control reviews, system exception reports, workflow approval steps that capture rejected items. If QC isn’t currently capturing pre-correction error rates, the sampling methodology is the first thing to build.
| Threshold | Criteria |
|---|---|
| Green | Error rate <2% across non-critical processes; <0.5% in regulatory reporting, payments, and compliance-sensitive processes |
| Amber | 2–5% in any non-critical process; or error rate trending upward three consecutive months; or any error rate >0.5% in regulatory or compliance-adjacent processes |
| Red | >5% in any process; or any error in a regulatory filing, consent order reporting, or regulatory threshold submission |
Owner: Quality control or operations. Risk and compliance own the threshold definitions for regulatory-adjacent processes, and should receive automatic alerts at the Amber threshold.
KRI 5: Rework Rate
What it measures: The percentage of completed items that require correction after initial processing — transactions voided and resubmitted, documents returned for revision, approvals reversed and reissued, reports sent back for correction.
Why it matters: Rework rate captures the downstream cost of process failures that error rate misses. A low first-pass error rate can coexist with a high rework rate if errors are caught late in the workflow — meaning the initial step got it wrong, the error survived several downstream handoffs, and correction now requires unwinding multiple steps.
High rework rates also create hidden capacity problems that compound over time. If 15% of your operations team’s daily capacity is spent correcting yesterday’s work rather than processing today’s volume, the backlog KRI will eventually turn amber even if nobody can explain why throughput declined.
Data source: Workflow systems with revision or return flags; quality review tracking; exception handling logs that distinguish first-pass completions from returns.
| Threshold | Criteria |
|---|---|
| Green | Rework rate <5% across all processes |
| Amber | 5–15%; or rework concentrated in a single process or business line above 10% |
| Red | >15%; or rework rate growing more than 20% quarter-over-quarter without documented volume spike explanation |
Owner: Operations lead with monthly trend analysis. Rework rates that drift gradually are harder to catch than sudden spikes — trend tracking matters more than point-in-time snapshots.
KRI 6: Handoff Failure Rate
What it measures: The percentage of process handoffs — from one team to another, one system to another, or one phase of a workflow to the next — that require remediation because the receiving party cannot proceed with what was delivered.
Why it matters: Most process failures don’t happen within a single team. They happen at the boundary. A loan operations team delivers a package to underwriting with missing documents. An exception queue is transferred between shifts with no status notes. A regulatory report prepared by operations is handed to compliance with errors that surface at the filing deadline.
Handoff failures are often invisible because neither party has a KRI for it. The sending team considers the job complete when the handoff happens. The receiving team treats the remediation as their own workload. Neither the volume nor the pattern makes it into risk reporting.
Data source: Rejection logs at receiving teams; return-to-sender records in workflow tools; escalation tickets that reference the prior team’s output as the root cause.
| Threshold | Criteria |
|---|---|
| Green | <3% of handoffs require remediation on receipt |
| Amber | 3–10%; or handoff failure rate concentrated in a specific sending team above 5% |
| Red | >10%; or any handoff failure affecting a regulatory reporting or customer-facing process |
Owner: Co-owned across the boundary — the sending team tracks outbound rejection rates, the receiving team tracks inbound return rates, and both are compared at least quarterly to surface chronic versus episodic patterns.
Summary Table
| KRI | Primary Signal | Data Source | Review Cadence |
|---|---|---|---|
| Backlog Aging by Process Type | Queue accumulation and unresolved items | Ticketing / workflow systems | Monthly |
| Manual Workaround Rate | Undocumented process bypasses | Walkthroughs, ops surveys | Quarterly |
| SLA Breach Rate | Timing failures, regulatory exposure | SLA trackers, ops dashboards | Monthly |
| Error Rate by Process Type | Quality degradation, compliance exposure | QC reviews, exception reports | Monthly |
| Rework Rate | Hidden capacity cost from upstream errors | Workflow revision logs | Monthly |
| Handoff Failure Rate | Cross-team breakage at process boundaries | Rejection and return logs | Monthly |
How Process KRIs Connect to the Rest of Your Risk Program
Process risk doesn’t operate in isolation. A sustained high rework rate in loan operations generates customer complaints — which surface in complaint management metrics and eventually in control testing KRI exception rates. A rising SLA breach rate in regulatory reporting generates issues that need to be tracked, assigned owners, and remediated — and those issues show up in issue management KRI aging.
The co-movement matters for escalation decisions. Process risk KRIs turning amber at the same time your operational risk KRI program is showing elevated exception rates in adjacent areas isn’t two separate data points — it’s a pattern worth escalating as a connected risk.
The OSFI Operational Risk Management and Resilience Guideline captures this connection directly: effective operational risk management requires identifying risk across process, people, systems, and external events as an interconnected system, not a collection of isolated metrics.
What Examiners See When These KRIs Are Missing
OCC examination teams reviewing operational risk management programs look for evidence that management knows when key processes are under stress — not just when losses have occurred. The absence of process risk KRIs doesn’t prove the risk isn’t being monitored. But it means the institution cannot demonstrate proactive monitoring to an examiner who asks.
Common patterns examiners identify at institutions without process risk KRIs:
Backlogs that management learns about when they become crises. The reconciliation team knew the queue was growing for three months. Risk management found out when an unresolved item hit the P&L and needed to be escalated to the board.
Manual workarounds that become permanent infrastructure. A system gap from a 2022 platform migration is still being handled by one analyst in a spreadsheet. Nobody told risk. Nobody built a compensating control around it. The analyst is the control.
Error rates that look acceptable until volume changes. A 1% error rate at 50,000 transactions per month becomes 500 errors per day at 50 million — and nobody updated the threshold, the staffing, or the QC process to account for the change.
Building these six KRIs doesn’t require new tooling. It requires identifying where your operational data lives, assigning a metric owner to each process, and defining a threshold that reflects genuine stress rather than business-as-usual variation.
So What?
Process risk is where operational losses start — in the backlog nobody escalated, the workaround that became policy, the handoff that failed quietly for months before anyone outside the ops team knew. The KRIs above don’t prevent process failure. They surface it early enough that management can act before it becomes a loss event, an audit finding, or an examiner’s observation.
If you can’t answer the following questions from existing data today, that’s your starting point:
- Which processes have items older than 30 days in queue right now?
- What percentage of this week’s transaction volume involved a step that wasn’t in the designed process?
- Which process had the highest error rate last month — and does anyone outside operations know?
The KRI Library (132 Key Risk Indicators) includes pre-built operational process KRIs with threshold ranges calibrated for financial services programs, alongside indicators across compliance, cyber, vendor, and BSA/AML risk domains.
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
KRI Library (132 Key Risk Indicators)
132 KRIs with thresholds, data sources, and escalation triggers pre-built for financial services.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What is a process risk KRI?
Why do regulators care about process backlogs and manual workarounds?
What is the difference between error rate and rework rate?
How should I calibrate thresholds for process risk KRIs?
Who should own process risk KRIs?
Which processes should I prioritize for process risk KRI coverage first?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
KRI Library (132 Key Risk Indicators)
132 KRIs with thresholds, data sources, and escalation triggers pre-built for financial services.
◆ Keep reading
Related posts.
Operational Risk
Risk Assessment Template in Excel: Build the Evidence Trail, Not Just the Heat Map
Build a risk assessment template in Excel that preserves evidence, challenge, approvals, and score history—not just a polished heat map.
Jul 23, 2026
Operational Risk
FedNow's Network Intelligence API Launched in April 2026. Your Fraud Risk Program Probably Hasn't Caught Up.
On April 28, 2026, the Federal Reserve made pre-payment network-level fraud intelligence available to every FedNow participant. The data — receiver account behavioral trends derived from system-wide FedNow activity — is available before a transaction is approved. Most institutions haven't updated their fraud policies, controls, or KRIs to account for what this changes.
Jul 21, 2026
Operational Risk
3,383 Incidents Later: What DORA's First ICT Data Reveals About Your Operational Risk Program
The ESAs published their first DORA ICT incident report in June 2026 — 3,383 major incidents, nearly one-third from third-party failures, only 10% cyber-related. Here's what the data means for your operational risk program.
Jul 16, 2026