Skip to content

Scenario-Based Retrieval Test

Retrieval test — instantiates Encoding–Retrieval Context Alignment

Judges recall by staging realistic scenarios that supply the authentic retrieval cues, then scoring whether the right knowledge surfaces — measuring readiness under representative demand, not bare recognition.

A quiz can tell you someone has the knowledge; it cannot tell you it will come back when the real situation asks for it in its own terms. Scenario-Based Retrieval Test closes that gap by measuring recall under representative demand: it stages a realistic situation that supplies the cues real use would supply — and withholds the scaffolding of a quiz — then scores whether the right knowledge actually surfaces and acts. Its defining move is that the test items reproduce the retrieval context, so a pass predicts field performance rather than the ability to recognize a correct answer among options. It is a measurement, not practice: its product is a signal about who is ready and where the gaps are, not a rehearsal.

Example

A company wants to know whether its field sales staff will actually apply the anti-bribery policy, not just pass the policy quiz. So instead of "which of the following is prohibited," it runs a situational-judgment-style scenario: a distributor offers to "expedite" a stalled permit for a modest cash gift, the quarter closes tomorrow, and your number is short. The rep has to produce the move. The results are revealing — several people who scored full marks on the written policy freeze or rationalize in the scenario, because recognition of a rule is not the same as retrieving it when a live, pressured situation supplies competing cues. The scored outcome flags exactly those reps for follow-up, and flags which part of the policy fails to surface under demand. The test measured the thing that matters: retrieval when the situation, not the exam, sets the cues.

How it works

What distinguishes it from a recognition quiz is that it engineers the retrieval cues of real use into the items and strips out the training scaffolding, then reads the outcome. It models the demand — what situation will call for this knowledge, and with what cues — builds scenarios that supply those cues (and, pointedly, the absence of the classroom's prompts), and scores each attempt into a signal: pass, fail, and the failure mode. A recognition probe hands you the answer to identify; an interference test crowds you with competitors; this test's distinctive job is representativeness — making the demand look like the world so the score means something about the world.

Tuning parameters

  • Scenario representativeness — how closely each item's cues match real use. Higher representativeness makes the score predictive but costs more to author and to grade.
  • Cue stripping — how much training scaffolding (headings, hints, prompts) is removed. More stripping tests genuine retrieval but can under-credit partial knowledge that a nudge would surface.
  • Scoring granularity — pass/fail versus a graded diagnosis of how it failed. Finer scoring localizes the gap for repair but costs rater effort and consistency.
  • Consequence realism — how consequential the scenario feels. More realism recruits the real internal state and sharpens the signal, but raises test anxiety that can distort it.
  • Breadth vs. depth — many short scenarios or a few deep ones. Breadth samples more of the demand space; depth catches subtle, situation-specific failures.

When it helps, and when it misleads

Its strength is that it separates I recognize it from I can retrieve and act on it under the cues of real use — catching context-bound fluency before it fails in the field, and telling you not just whether but where recall breaks.

Its failure mode is that a test is only as valid as its scenarios are representative: items that look realistic but supply subtly different cues measure the wrong thing and reassure (or condemn) falsely. The classic misuse is teaching to the test — drilling the exact scenarios — which re-binds recall to the test's own cues and quietly reintroduces the context-dependence the test exists to detect. The discipline that guards against this is representative design[^repdesign]: sample the scenario space to match the real distribution of demands, rotate items so they cannot be memorized, and trust the score only to the degree the scenarios stand in for actual use.

How it implements the components

Scenario-Based Retrieval Test fills the measure-under-demand side of the archetype — the parts a diagnostic instrument can produce:

  • representative_transfer_test — it builds items whose cues represent the conditions of real use, so a pass measures transfer rather than recognition.
  • retrieval_outcome_signal — it scores each attempt into a pass/fail-plus-failure-mode signal that downstream mechanisms can act on.

It does not build the practice environment those scenarios might live in — that is Representative-Environment Simulation's job — nor localize availability-versus-accessibility (Free-Recall-Then-Recognition Probe) or stress competing traces (Interleaved Competitor Retrieval Test).

References

Representative design (Brunswik) — a test predicts real-world performance only to the degree its conditions are sampled to represent the conditions of actual use. A scenario test's validity rests entirely on how representative its scenarios are of the moment they stand in for.