Feature Business Continuity
Tabletop Exercise Evaluation Rubric: Ratings, Assignments, and Observation Notes
Build a tabletop exercise evaluation rubric with evaluator assignments, evidence-based ratings, observation notes, calibration, and AAR handoff.
Table of Contents
TL;DR
- Assign evaluators to objectives and evidence streams, not merely to seats in the room. Every priority objective needs named coverage and a backup.
- Rate observable performance against prewritten criteria. Confidence, airtime, and eventual consensus are context—not proof that a capability worked.
- Write notes as fact + reference + impact: what happened, what plan or target applied, and why the difference mattered.
- Calibrate ratings before the exercise and reconcile evidence afterward. Do not average contradictory judgments into a reassuring score.
The most dangerous tabletop participant is the evaluator with a blank notebook and no criteria.
They will capture who spoke, which comments sounded smart, and whether the room felt organized. Two hours later, the exercise gets a “successful” rating because everyone participated. Meanwhile, nobody recorded that activation authority was unclear, the payment-processor contact was stale, or the workaround exceeded its capacity before anyone approved customer prioritization.
A tabletop exercise evaluation rubric fixes that by defining what evaluators should observe, who covers each objective, what evidence supports a rating, and how observations become after-action findings. The rubric is not a generic scorecard. It is the bridge between exercise design and a defensible conclusion.
This article starts where the tabletop MSEL guide stops. The MSEL controls events and expected actions. The rubric evaluates performance across those events. For facilitation and probing, use the tabletop facilitation guide. For the downstream report, use the BCP after-action report guide.
Start With HSEEP’s Evaluation Logic—Then Keep It Proportionate
FEMA’s 2020 Homeland Security Exercise and Evaluation Program doctrine describes Exercise Evaluation Guides, or EEGs, as tools aligned to objectives, capabilities, capability targets, and critical tasks. HSEEP says evaluation planning includes selecting team requirements, developing documentation and methodology, and preparing the EEGs.
Its observation guidance is equally useful for a financial-services tabletop. Evaluators collect facts about plan activation, roles and authorities, decisions, and information sharing. Data collection should create a fact-based record of actions, decisions, and outcomes instead of relying on assumptions.
HSEEP is exercise doctrine, not a bank regulation. The FEMA HSEEP program page describes the broader methodology and resources. The FFIEC Business Continuity Management tabletop section is financial-services examination guidance; it does not require this article’s exact rubric or rating labels.
Use the logic, not the bureaucracy. A 90-minute payments-outage tabletop does not need a government-scale evaluation organization. It does need explicit objectives, observable decisions, retained evidence, and a controlled way to reach conclusions.
Build Evaluator Assignments Around Coverage
Assigning one evaluator to “Operations” and another to “Compliance” sounds organized, but it leaves gaps whenever an objective crosses functions. A customer-communication objective may involve Operations identifying impact, Legal constraining claims, Compliance assessing notices, and Communications approving language.
Use an evaluator coverage matrix before exercise day:
| Objective | What must be observed | Primary evaluator | Backup / second view | Evidence stream |
|---|---|---|---|---|
| OBJ-1: Classify and activate | Severity assessment, activation decision, authority, timestamp | Lead evaluator | Governance evaluator | Decision log, plan section, meeting record |
| OBJ-2: Protect customers during constrained operations | Prioritization criteria, control limits, approver, unresolved exposure | Operations evaluator | Compliance evaluator | Workaround worksheet, queue data, approval |
| OBJ-3: Communicate accurately | Audience, known facts, uncertainty, approval, next update | Communications evaluator | Legal/Compliance evaluator | Draft message, approval trail |
| OBJ-4: Recover under control | Reconciliation, validation, exception ownership, exit criteria | Technology evaluator | Operations evaluator | Recovery checklist, reconciliation plan |
The table is a realistic hypothetical, not a staffing mandate.
The lead evaluator owns cross-objective consistency and final synthesis. Objective evaluators capture detailed facts. A note taker can support the record, but should not become the only person who knows what happened. If the tabletop uses breakout rooms, assign coverage to each room and define how notes return to the lead.
Before play, each evaluator should receive:
- exercise scope, objectives, and scenario boundaries;
- the applicable plan, procedure, authority matrix, and approved targets;
- assigned objectives and injects;
- expected observable actions and evidence;
- rating definitions and examples;
- note format and evidence-naming convention; and
- escalation instructions for safety, real-world incidents, or an unobservable objective.
HSEEP specifically calls for evaluator assignments, relevant expertise, and training or instruction before the exercise. Where possible, give each objective to someone who understands the process but is not responsible for defending its design.
Use a Rating Scale That Separates Performance From Evidence
A rating should answer: did participants demonstrate the objective, and can the evaluation team support that conclusion?
The following four-level scale is an optional internal design. It is not prescribed by FEMA, FFIEC, NIST, or ISO.
| Rating | Decision rule | Minimum documentation |
|---|---|---|
| Demonstrated | Critical actions occurred within the exercise criteria; authority and outputs were clear; retained evidence supports the result | Facts, timestamps, plan/target reference, artifact IDs |
| Partially demonstrated | The core objective was achieved, but one or more material steps, authorities, dependencies, or outputs were late, incomplete, or unclear | Achieved elements, gaps, impact, evidence |
| Not demonstrated | Participants did not complete the critical action, used an unsupported path, or produced an outcome that did not meet the objective | What occurred, expected criterion, consequence, evidence |
| Not evaluated | The exercise did not create a fair opportunity to observe the objective, or evidence is insufficient to rate it | Reason, design/coverage gap, recommended retest |
“Not evaluated” protects the integrity of the rubric. If an inject was skipped, a breakout room had no evaluator, or the scenario never reached recovery, do not convert missing evidence into a pass or fail.
Avoid a single overall percentage. Averaging “Demonstrated” activation with “Not demonstrated” recovery can produce a middle score that hides the exact capability likely to hurt customers. Report objective-level ratings and identify any critical task that caps the conclusion.
A workable decision rule might say: the overall exercise cannot be rated demonstrated if any priority objective is not demonstrated or not evaluated. That is an internal governance choice. Approve it before the session and tailor it to exercise purpose; do not present it as an industry requirement.
Write Observation Notes That Survive Challenge
Useful evaluator notes have three parts:
- Fact: What was said, decided, produced, or omitted, with a time or inject reference.
- Reference: What plan provision, target, authority, or expected action applied.
- Impact: Why the alignment or difference mattered to the objective.
Call it the FRI note if your team needs a mnemonic.
Weak note
Team was confused about escalation.
Stronger note
Fact: At 09:24, after INJ-03, participants named three possible incident approvers and paused the activation decision for 11 minutes; no participant opened the delegation matrix. Reference: BCP §3.2 assigns activation to the Incident Lead or named delegate. Impact: The delay left workaround approvals without confirmed authority during the simulated customer-impact period. Evidence: OBS-OBJ1-04 and exercise recording 00:24–00:35.
That example is hypothetical. It separates what happened from the evaluator’s analysis and points the reviewer to evidence.
For every priority observation, capture:
- objective and inject IDs;
- simulated and actual time;
- role involved, without turning the record into a personnel scorecard;
- action or decision;
- information available at that moment;
- applicable plan, target, or expected action;
- artifact or note reference;
- result or consequence; and
- initial strength, gap, or follow-up question.
HSEEP calls for observation in a non-attribution environment. Unless an exercise specifically tests an individual qualification, evaluate the capability and process. “The incident-lead role did not consult the delegation matrix” is more useful than naming an employee as the failure.
Do not write the after-action recommendation during live play. Evaluators should capture facts first. Early solution-writing creates confirmation bias: once the note says “training issue,” the evaluator may stop looking for the outdated authority matrix that caused the confusion.
Distinguish a Real Gap From an Exercise Artifact
A tabletop can fail because the capability is weak, the plan is weak, or the exercise design is weak. The rubric should keep those apart.
| What happened | Likely classification | Evaluation response |
|---|---|---|
| Participants had the needed facts but could not identify the authorized decision-maker | Capability or plan gap | Rate against objective; inspect authority source |
| The scenario withheld information that would exist in a real incident | Design limitation | Mark affected element not evaluated or qualify the rating |
| Controller coached participants repeatedly until they reached the expected answer | Performance gap masked by control intervention | Record intervention; do not treat final answer as independent demonstration |
| Current plan supported a different valid decision than the expected action | Evaluation-criteria defect or plan/design conflict | Preserve evidence; escalate for calibration before rating |
| Evaluator missed the breakout where the decision occurred | Coverage failure | Seek independent artifacts; otherwise mark not evaluated |
This is where ownership gets messy. Exercise designers naturally defend the scenario. Process owners defend their performance. Evaluators defend their notes. The lead evaluator’s job is not to split the difference. It is to reconstruct what happened and classify the source of the gap using evidence.
Calibrate Before Play, Reconcile After Play
Run a short evaluator calibration using one hypothetical observation. Ask every evaluator to rate it and explain which evidence drove the rating. Differences before exercise day are cheap to fix.
Calibration should settle questions such as:
- Does a correct decision without a documented approver count as partial or not demonstrated?
- If participants describe an artifact but do not create it, what does the objective require?
- Does controller prompting affect the rating?
- Which tasks are critical enough to cap the objective?
- How are valid alternative actions evaluated?
- When does missing evidence become “not evaluated” rather than “not demonstrated”?
After play, hold an evaluator-only reconciliation before the general hot wash changes the narrative. Review each objective in this order:
- planned criterion;
- timestamped observations;
- retained artifacts;
- conflicting evidence;
- provisional rating and rationale; and
- candidate strength or area for improvement.
If two evaluators disagree, compare evidence—not adjectives. One may have heard the decision while another saw that it was never entered into the log. Both facts can be true, and together they may support “partially demonstrated.”
Do not average ratings. If the conflict cannot be resolved, preserve the limitation and request an artifact or targeted retest. An inconclusive rating is more defensible than fake precision.
Hand the Rubric Into the AAR Without Losing Traceability
HSEEP’s evaluation chapter says evaluators compare actual performance with objectives, identify strengths and areas for improvement, and use analysis such as data synthesis, event reconstruction, trend analysis, and root-cause analysis. It describes strong AAR observations as a clear issue statement, analysis, and impact.
Build the handoff at the row level:
objective → inject → observation → evidence → rating → AAR observation → corrective action
A handoff table can look like this:
| Objective | Rating | Observation IDs | AAR disposition | Owner action |
|---|---|---|---|---|
| OBJ-1 | Partially demonstrated | OBS-01, OBS-04 | Area for improvement: activation authority not operationalized | BCM owner updates delegation job aid; Incident Management validates in retest |
| OBJ-2 | Demonstrated | OBS-06, OBS-08 | Strength: prioritization criteria applied with approval | Preserve artifact as worked example |
| OBJ-3 | Not evaluated | OBS-11 | Exercise-design limitation: communications inject skipped | Add objective to next targeted drill |
Again, those records are hypothetical.
The after-action meeting should validate findings and corrective actions, not rewrite observed facts to make the room more comfortable. HSEEP says the meeting seeks consensus on strengths, areas for improvement, corrective actions, deadlines, and owners. If stakeholders dispute an observation, append evidence and resolve the analysis. Do not delete the timestamped note.
NIST SP 800-84, Guide to Test, Training, and Exercise Programs for IT Plans and Capabilities offers a useful adjacent control: exercise and test findings, observations, and enhancement considerations should feed an after-action report. It is voluntary technical guidance outside its federal scope, not a bank-specific rating requirement.
So What?
Before the next tabletop, do not start by printing blank observation forms. Build a one-page coverage matrix and one-page rating guide.
Name the evaluator for every objective. Define the critical decisions and evidence. Calibrate one example. Require fact-reference-impact notes. Then reconcile the record before the hot wash turns everyone’s memory into a cleaner story than what actually happened.
A good rubric will sometimes produce “not evaluated.” That is not an embarrassing result. It is evidence that the evaluation team can distinguish capability performance from missing coverage—and is willing to schedule the test that still needs to happen.
The Business Continuity & Disaster Recovery (BCP/DR) Kit includes a facilitator-ready tabletop kit, findings template, test log, and action tracker; add the objective-level rating and evidence fields above to fit your exercise.
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What should a tabletop exercise evaluation rubric include?
Does FEMA require a four-level rating scale for tabletop exercises?
Should evaluators rate individual participants?
How many evaluators does a tabletop need?
How should conflicting evaluator ratings be resolved?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Business Continuity & Disaster Recovery (BCP/DR) Kit
BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.
◆ Keep reading
Related posts.
Business Continuity
ISO 22301 Clause 9.3 Management Review: Agenda, Inputs, Decisions, and Evidence
Build an ISO 22301 management review decision pack that maps Clause 9.3 inputs to evidence, decisions, owners, and follow-up.
Aug 21, 2026
Business Continuity
BIA Quality Assurance: Challenge Outliers and Calibrate Scores Across Business Units
Run business impact analysis quality assurance with outlier tests, challenge notes, score calibration, capability checks, and approval evidence.
Aug 20, 2026
Business Continuity
Operational Resilience Critical-Service Dependency Crosswalk: Reconcile Processes, Systems, Vendors, Facilities, and People
A critical-service crosswalk reconciles your process, system, vendor, facility, and people inventories under a shared service ID—so your resilience program has one source of truth instead of five competing lists. Here's how to build it.
Aug 19, 2026