Skip to content
RiskTemplates · The Daily Brief Friday, August 21, 2026
Wire SEC's Tricolor Fraud Case: The Double-Pledging Controls Lenders Missed AUG 20

Feature Business Continuity

Tabletop Exercise Evaluation Rubric: Ratings, Assignments, and Observation Notes

Build a tabletop exercise evaluation rubric with evaluator assignments, evidence-based ratings, observation notes, calibration, and AAR handoff.

Table of Contents

TL;DR

  • Assign evaluators to objectives and evidence streams, not merely to seats in the room. Every priority objective needs named coverage and a backup.
  • Rate observable performance against prewritten criteria. Confidence, airtime, and eventual consensus are context—not proof that a capability worked.
  • Write notes as fact + reference + impact: what happened, what plan or target applied, and why the difference mattered.
  • Calibrate ratings before the exercise and reconcile evidence afterward. Do not average contradictory judgments into a reassuring score.

The most dangerous tabletop participant is the evaluator with a blank notebook and no criteria.

They will capture who spoke, which comments sounded smart, and whether the room felt organized. Two hours later, the exercise gets a “successful” rating because everyone participated. Meanwhile, nobody recorded that activation authority was unclear, the payment-processor contact was stale, or the workaround exceeded its capacity before anyone approved customer prioritization.

A tabletop exercise evaluation rubric fixes that by defining what evaluators should observe, who covers each objective, what evidence supports a rating, and how observations become after-action findings. The rubric is not a generic scorecard. It is the bridge between exercise design and a defensible conclusion.

This article starts where the tabletop MSEL guide stops. The MSEL controls events and expected actions. The rubric evaluates performance across those events. For facilitation and probing, use the tabletop facilitation guide. For the downstream report, use the BCP after-action report guide.

Start With HSEEP’s Evaluation Logic—Then Keep It Proportionate

FEMA’s 2020 Homeland Security Exercise and Evaluation Program doctrine describes Exercise Evaluation Guides, or EEGs, as tools aligned to objectives, capabilities, capability targets, and critical tasks. HSEEP says evaluation planning includes selecting team requirements, developing documentation and methodology, and preparing the EEGs.

Its observation guidance is equally useful for a financial-services tabletop. Evaluators collect facts about plan activation, roles and authorities, decisions, and information sharing. Data collection should create a fact-based record of actions, decisions, and outcomes instead of relying on assumptions.

HSEEP is exercise doctrine, not a bank regulation. The FEMA HSEEP program page describes the broader methodology and resources. The FFIEC Business Continuity Management tabletop section is financial-services examination guidance; it does not require this article’s exact rubric or rating labels.

Use the logic, not the bureaucracy. A 90-minute payments-outage tabletop does not need a government-scale evaluation organization. It does need explicit objectives, observable decisions, retained evidence, and a controlled way to reach conclusions.

Build Evaluator Assignments Around Coverage

Assigning one evaluator to “Operations” and another to “Compliance” sounds organized, but it leaves gaps whenever an objective crosses functions. A customer-communication objective may involve Operations identifying impact, Legal constraining claims, Compliance assessing notices, and Communications approving language.

Use an evaluator coverage matrix before exercise day:

ObjectiveWhat must be observedPrimary evaluatorBackup / second viewEvidence stream
OBJ-1: Classify and activateSeverity assessment, activation decision, authority, timestampLead evaluatorGovernance evaluatorDecision log, plan section, meeting record
OBJ-2: Protect customers during constrained operationsPrioritization criteria, control limits, approver, unresolved exposureOperations evaluatorCompliance evaluatorWorkaround worksheet, queue data, approval
OBJ-3: Communicate accuratelyAudience, known facts, uncertainty, approval, next updateCommunications evaluatorLegal/Compliance evaluatorDraft message, approval trail
OBJ-4: Recover under controlReconciliation, validation, exception ownership, exit criteriaTechnology evaluatorOperations evaluatorRecovery checklist, reconciliation plan

The table is a realistic hypothetical, not a staffing mandate.

The lead evaluator owns cross-objective consistency and final synthesis. Objective evaluators capture detailed facts. A note taker can support the record, but should not become the only person who knows what happened. If the tabletop uses breakout rooms, assign coverage to each room and define how notes return to the lead.

Before play, each evaluator should receive:

  • exercise scope, objectives, and scenario boundaries;
  • the applicable plan, procedure, authority matrix, and approved targets;
  • assigned objectives and injects;
  • expected observable actions and evidence;
  • rating definitions and examples;
  • note format and evidence-naming convention; and
  • escalation instructions for safety, real-world incidents, or an unobservable objective.

HSEEP specifically calls for evaluator assignments, relevant expertise, and training or instruction before the exercise. Where possible, give each objective to someone who understands the process but is not responsible for defending its design.

Use a Rating Scale That Separates Performance From Evidence

A rating should answer: did participants demonstrate the objective, and can the evaluation team support that conclusion?

The following four-level scale is an optional internal design. It is not prescribed by FEMA, FFIEC, NIST, or ISO.

RatingDecision ruleMinimum documentation
DemonstratedCritical actions occurred within the exercise criteria; authority and outputs were clear; retained evidence supports the resultFacts, timestamps, plan/target reference, artifact IDs
Partially demonstratedThe core objective was achieved, but one or more material steps, authorities, dependencies, or outputs were late, incomplete, or unclearAchieved elements, gaps, impact, evidence
Not demonstratedParticipants did not complete the critical action, used an unsupported path, or produced an outcome that did not meet the objectiveWhat occurred, expected criterion, consequence, evidence
Not evaluatedThe exercise did not create a fair opportunity to observe the objective, or evidence is insufficient to rate itReason, design/coverage gap, recommended retest

“Not evaluated” protects the integrity of the rubric. If an inject was skipped, a breakout room had no evaluator, or the scenario never reached recovery, do not convert missing evidence into a pass or fail.

Avoid a single overall percentage. Averaging “Demonstrated” activation with “Not demonstrated” recovery can produce a middle score that hides the exact capability likely to hurt customers. Report objective-level ratings and identify any critical task that caps the conclusion.

A workable decision rule might say: the overall exercise cannot be rated demonstrated if any priority objective is not demonstrated or not evaluated. That is an internal governance choice. Approve it before the session and tailor it to exercise purpose; do not present it as an industry requirement.

Write Observation Notes That Survive Challenge

Useful evaluator notes have three parts:

  1. Fact: What was said, decided, produced, or omitted, with a time or inject reference.
  2. Reference: What plan provision, target, authority, or expected action applied.
  3. Impact: Why the alignment or difference mattered to the objective.

Call it the FRI note if your team needs a mnemonic.

Weak note

Team was confused about escalation.

Stronger note

Fact: At 09:24, after INJ-03, participants named three possible incident approvers and paused the activation decision for 11 minutes; no participant opened the delegation matrix. Reference: BCP §3.2 assigns activation to the Incident Lead or named delegate. Impact: The delay left workaround approvals without confirmed authority during the simulated customer-impact period. Evidence: OBS-OBJ1-04 and exercise recording 00:24–00:35.

That example is hypothetical. It separates what happened from the evaluator’s analysis and points the reviewer to evidence.

For every priority observation, capture:

  • objective and inject IDs;
  • simulated and actual time;
  • role involved, without turning the record into a personnel scorecard;
  • action or decision;
  • information available at that moment;
  • applicable plan, target, or expected action;
  • artifact or note reference;
  • result or consequence; and
  • initial strength, gap, or follow-up question.

HSEEP calls for observation in a non-attribution environment. Unless an exercise specifically tests an individual qualification, evaluate the capability and process. “The incident-lead role did not consult the delegation matrix” is more useful than naming an employee as the failure.

Do not write the after-action recommendation during live play. Evaluators should capture facts first. Early solution-writing creates confirmation bias: once the note says “training issue,” the evaluator may stop looking for the outdated authority matrix that caused the confusion.

Distinguish a Real Gap From an Exercise Artifact

A tabletop can fail because the capability is weak, the plan is weak, or the exercise design is weak. The rubric should keep those apart.

What happenedLikely classificationEvaluation response
Participants had the needed facts but could not identify the authorized decision-makerCapability or plan gapRate against objective; inspect authority source
The scenario withheld information that would exist in a real incidentDesign limitationMark affected element not evaluated or qualify the rating
Controller coached participants repeatedly until they reached the expected answerPerformance gap masked by control interventionRecord intervention; do not treat final answer as independent demonstration
Current plan supported a different valid decision than the expected actionEvaluation-criteria defect or plan/design conflictPreserve evidence; escalate for calibration before rating
Evaluator missed the breakout where the decision occurredCoverage failureSeek independent artifacts; otherwise mark not evaluated

This is where ownership gets messy. Exercise designers naturally defend the scenario. Process owners defend their performance. Evaluators defend their notes. The lead evaluator’s job is not to split the difference. It is to reconstruct what happened and classify the source of the gap using evidence.

Calibrate Before Play, Reconcile After Play

Run a short evaluator calibration using one hypothetical observation. Ask every evaluator to rate it and explain which evidence drove the rating. Differences before exercise day are cheap to fix.

Calibration should settle questions such as:

  • Does a correct decision without a documented approver count as partial or not demonstrated?
  • If participants describe an artifact but do not create it, what does the objective require?
  • Does controller prompting affect the rating?
  • Which tasks are critical enough to cap the objective?
  • How are valid alternative actions evaluated?
  • When does missing evidence become “not evaluated” rather than “not demonstrated”?

After play, hold an evaluator-only reconciliation before the general hot wash changes the narrative. Review each objective in this order:

  1. planned criterion;
  2. timestamped observations;
  3. retained artifacts;
  4. conflicting evidence;
  5. provisional rating and rationale; and
  6. candidate strength or area for improvement.

If two evaluators disagree, compare evidence—not adjectives. One may have heard the decision while another saw that it was never entered into the log. Both facts can be true, and together they may support “partially demonstrated.”

Do not average ratings. If the conflict cannot be resolved, preserve the limitation and request an artifact or targeted retest. An inconclusive rating is more defensible than fake precision.

Hand the Rubric Into the AAR Without Losing Traceability

HSEEP’s evaluation chapter says evaluators compare actual performance with objectives, identify strengths and areas for improvement, and use analysis such as data synthesis, event reconstruction, trend analysis, and root-cause analysis. It describes strong AAR observations as a clear issue statement, analysis, and impact.

Build the handoff at the row level:

objective → inject → observation → evidence → rating → AAR observation → corrective action

A handoff table can look like this:

ObjectiveRatingObservation IDsAAR dispositionOwner action
OBJ-1Partially demonstratedOBS-01, OBS-04Area for improvement: activation authority not operationalizedBCM owner updates delegation job aid; Incident Management validates in retest
OBJ-2DemonstratedOBS-06, OBS-08Strength: prioritization criteria applied with approvalPreserve artifact as worked example
OBJ-3Not evaluatedOBS-11Exercise-design limitation: communications inject skippedAdd objective to next targeted drill

Again, those records are hypothetical.

The after-action meeting should validate findings and corrective actions, not rewrite observed facts to make the room more comfortable. HSEEP says the meeting seeks consensus on strengths, areas for improvement, corrective actions, deadlines, and owners. If stakeholders dispute an observation, append evidence and resolve the analysis. Do not delete the timestamped note.

NIST SP 800-84, Guide to Test, Training, and Exercise Programs for IT Plans and Capabilities offers a useful adjacent control: exercise and test findings, observations, and enhancement considerations should feed an after-action report. It is voluntary technical guidance outside its federal scope, not a bank-specific rating requirement.

So What?

Before the next tabletop, do not start by printing blank observation forms. Build a one-page coverage matrix and one-page rating guide.

Name the evaluator for every objective. Define the critical decisions and evidence. Calibrate one example. Require fact-reference-impact notes. Then reconcile the record before the hot wash turns everyone’s memory into a cleaner story than what actually happened.

A good rubric will sometimes produce “not evaluated.” That is not an embarrassing result. It is evidence that the evaluation team can distinguish capability performance from missing coverage—and is willing to schedule the test that still needs to happen.

The Business Continuity & Disaster Recovery (BCP/DR) Kit includes a facilitator-ready tabletop kit, findings template, test log, and action tracker; add the objective-level rating and evidence fields above to fit your exercise.

◆ Need the working template?

Start with the source guide.

These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.

◆ Immaterial Findings · Weekly

Sharp risk & compliance insights. No fluff.

◆ FAQ

Frequently asked questions.

What should a tabletop exercise evaluation rubric include?
Include exercise objectives, observable tasks or decisions, expected evidence, evaluator assignments, a defined rating scale, timestamped observation notes, evidence references, gaps and strengths, and a clear handoff into the after-action report and corrective-action process.
Does FEMA require a four-level rating scale for tabletop exercises?
No. FEMA HSEEP provides an evaluation method built around objectives, capability targets, critical tasks, observation, evidence, and after-action analysis. The four-level scale in this article is an optional internal rubric for financial-services tabletops, not a FEMA or FFIEC mandate.
Should evaluators rate individual participants?
Usually no. Rate the capability, objective, process, or decision path unless the exercise specifically assesses an individual qualification. HSEEP calls for observation in a non-attribution environment; person-focused scoring can suppress useful discussion and distract from plan, authority, dependency, and control failures.
How many evaluators does a tabletop need?
There is no universal number. Assign enough coverage for the objectives, functional areas, locations, and evidence streams in scope. A focused tabletop may use one lead evaluator plus objective owners; a multi-room exercise may require evaluators at each venue or function. Document any coverage limitations.
How should conflicting evaluator ratings be resolved?
Compare the underlying observations, plan references, timestamps, and artifacts before debating labels. The lead evaluator should document the agreed rating and rationale. If evidence remains genuinely conflicting, preserve the disagreement or mark the objective inconclusive rather than averaging unsupported opinions.
Rebecca Leung

Author

Rebecca Leung

Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.

◆ Related framework

Business Continuity & Disaster Recovery (BCP/DR) Kit

BCP and DR templates with BIA, recovery procedures, and a standalone tabletop exercise kit.

Immaterial Findings · Newsletter

The brief, in your inbox.

Enforcement of the week, a framework breakdown, and the prompts that are actually worth running. Delivered to your inbox. Free.