Skip to content

Simulation Replay

Process — instantiates Tacit Knowledge Elicitation

Reconstructs a realistic case and plays it forward under controlled conditions, so what different practitioners notice and decide as it unfolds can be compared, and elicited knowledge tested against fresh eyes.

Real practice happens once and can't be paused, rewound, or given to two people to attempt independently. Simulation Replay manufactures that repeatability: it reconstructs a case as a scenario and plays it forward — in a simulator, a tabletop, or a recorded case run at controllable pace — so the same unfolding situation can be put in front of an expert, a novice, or a practitioner testing an elicited rule. Its defining move is controlled re-enactment for comparison and validation: because the scenario is fixed and repeatable, you can hold everything constant and vary only the person, watching what they notice, when, and how their reading changes as the case develops. That makes it the archetype's testing bench — the place where "an expert said this cue matters" becomes "here is a case where the cue appears; let's see who catches it and whether the rule holds."

Example

A grid operator builds a simulation of a specific cascading-failure event and replays it in the control-room trainer. First, the scenario is run with a veteran operator, whose actions and callouts are recorded second by second as the fault propagates. Then it's run with newer operators. The replays are laid side by side: the veteran shed load on one corridor forty seconds before the newcomers even flagged it as at-risk — a temporal judgment ("the moment this reading trends this way, it's already too late to wait") that only surfaces when the same clock runs for both.

The simulation then does double duty as a validation bench. A cue a debrief had elicited — "when frequency and this specific tie-line both dip together, pre-emptively island" — is tested by dropping fresh operators into a new scenario built to contain that signature: do they, given the elicited rule, actually make the veteran's call? Where they don't, the rule was incomplete or the scenario revealed a boundary. What the replay cannot do is generate the cue in the first place; it consumes what earlier elicitation surfaced and puts it to the test.

How it works

  • Reconstruct a realistic, repeatable case. Build a scenario faithful enough to evoke real judgment, but fixed so it can be run identically for many people.
  • Run it forward with the clock. Preserve the temporal unfolding — what is known when — so the mechanism can capture timing judgments (when to wait, when to act) that a static case review flattens.
  • Compare across practitioners. Put experts and novices through the same scenario and contrast what each notices and decides at each moment; the divergence maps tacit competence.
  • Test elicited knowledge on fresh eyes. Give a novice the articulated cue or rule and a new scenario containing it, and see whether they reproduce the expert judgment — validating (or breaking) the elicited claim in a realistic context.

Tuning parameters

  • Fidelity — how closely the simulation matches real conditions, stakes, and sensory load. Higher fidelity evokes truer judgment but costs more to build and can over-fit to the modeled case; lower fidelity is cheap but may not transfer.
  • Scripted vs. branching — a fixed replay or a scenario that responds to the practitioner's actions. Branching tests decision-making more realistically; scripted keeps the comparison across people clean.
  • Comparison design — who is run through it (experts, novices, or rule-armed testers) and against what baseline. This decides whether the session maps expertise, tests a novice, or validates a claim.
  • Pacing control — real-time, slowed, or pausable. Slowing lets observers catch fine reads; real-time preserves the pressure that produces authentic timing judgment.

When it helps, and when it misleads

Its strengths are repeatability and temporal insight: it recovers the when-to-act judgments that only appear as a case unfolds, it lets expert and novice be compared on identical ground, and it turns elicited knowledge into something testable — the closest this archetype comes to validating a claim before betting on it.

Its defining limitation is the fidelity gap: a simulation is not the real thing, and judgment that holds in the trainer can fail when real stakes, fatigue, and consequence return — so a rule "validated" only in simulation carries a hidden asterisk.[1] High-fidelity sims are expensive and can over-fit to the one reconstructed case, and practitioners may game the exercise, performing for the scenario rather than acting as they would for real. The classic misuse is treating simulation validation as proof of real-world transfer rather than as evidence with a boundary. The discipline that keeps it honest is to match fidelity to what's being tested, vary scenarios rather than validating on the single case a rule was derived from, and mark simulation-validated knowledge as provisionally transferable until it survives real conditions.

How it implements the components

Simulation Replay fills the comparison-and-validation components that require a repeatable, unfolding case:

  • expert_novice_comparison — the same scenario run for experts and novices makes the difference in what each notices, and when, directly observable.
  • contextual_validation — tests whether elicited knowledge actually works in a realistic reconstructed context, not just in the expert's telling.
  • novice_readback — a learner applies the elicited cue or rule to a fresh scenario, and their performance shows what really transferred.
  • replication_probe — checks whether a second practitioner, given the articulated knowledge, reproduces the expert judgment on a new case built to contain it.

It validates and compares but does not generate the raw material: the cues (cue_elicitation), the reasoning (decision_rationale_probe), and the incidents it reconstructs come from the eliciting mechanisms like Critical Incident Technique and Cognitive Interview; it also does not capture live real-world practice (expert_practice_observationShadowing Session).

Notes

Simulation Replay is where Think-Aloud Protocol is most safely run: when live narration would disrupt real work, a practitioner can narrate over a replayed scenario at controllable pace. The two compose — the replay supplies the repeatable case, the think-aloud supplies the concurrent inner stream — but they answer different questions (replay: who catches what, and does the rule hold; think-aloud: what is the attention stream), so they are kept as distinct mechanisms.

References

[1] The gap between performance in a simulation and performance in the real setting — a central concern of simulation fidelity in training research and of naturalistic decision-making, which stresses that real judgment is shaped by genuine stakes and consequence. It is why simulation-validated knowledge is treated above as provisionally, not proven, transferable.