Skip to content

Simulation or Case Test

Scenario-based test (tool) — instantiates Competence Calibration Feedback

Puts a person into a realistic simulated scenario or case and measures how they actually perform, generating high-fidelity evidence — including on unfamiliar situations — without real-world risk.

Simulation or Case Test manufactures the performance situation instead of waiting for a real one. Its defining move is to construct a realistic scenario — a full-mission simulator, a worked case — that can be dialed to demand exactly the competencies in question and, crucially, to throw unfamiliar variants that reveal whether skill actually generalizes or was merely memorized. Because the consequences are simulated, it can safely probe the high-stakes and rare situations real work rarely offers on demand. It generates observable evidence of what a person actually does under pressure; it does not itself judge that evidence or feed it back — that is for its siblings.

Example

An airline puts a first officer through a simulator session built around a scenario she has never trained: an engine failure at low altitude on approach, in weather, with a distracting secondary fault. This is deliberately not the textbook drill she can recite — it's a novel combination designed to test whether her competence transfers to a situation she hasn't rehearsed. The simulator captures everything: the sequence of her decisions, her callouts, how long she takes to stabilize, whether she keeps flying the aircraft while troubleshooting. The value isn't a pass/fail stamp; it's high-fidelity evidence of how she performs when the situation is unfamiliar and the pressure is real, gathered without a single passenger aboard. Full-mission scenario training of this kind is a long-standing practice precisely because it surfaces performance that calm, familiar testing hides.[1]

How it works

Its distinguishing feature is that it is a designed evidence source with two dials generic testing lacks. It can be tuned to the specific competencies under question, and — most importantly — its novelty can be raised to probe transfer: throw a variant far enough from the practiced case to distinguish real, generalizable skill from a memorized routine. Because it's simulated, it is repeatable, richly observable, and safe enough to run the failure and edge cases real work won't hand you. It produces the evidence; the calibration comparison and the feedback happen elsewhere.

Tuning parameters

  • Fidelity — how realistic the simulation is. Higher fidelity transfers better to real work but costs more; low fidelity is cheap but may test the wrong thing.
  • Novelty injection — how far the scenario departs from the familiar. More novelty probes transfer harder but risks testing luck instead of skill.
  • Scenario coverage — routine cases versus edge and failure cases. Edge cases reveal calibration under stress; routine ones confirm the baseline.
  • Observation richness — outcome-only versus capturing the process (decisions, timing, sequence). Process data reveals why performance was what it was.
  • Stakes realism — how much real pressure the scenario reproduces. Pressure surfaces the overconfidence that low-stakes testing lets hide.

When it helps, and when it misleads

Its strength is being the safe way to get evidence on high-stakes or rare situations, and the sharpest probe of whether competence transfers beyond the specific cases a person has practiced — the classic trap where familiarity is mistaken for mastery.[2]

It misleads when simulation performance fails to carry over: people can learn to fly the simulator rather than the aircraft, mastering the artifice instead of the job. A well-designed scenario can still miss the situations that actually bite in the real world, and passing a dramatic case test can breed exactly the overconfidence calibration is supposed to correct. The discipline is to vary and refresh scenarios so they can't be gamed, deliberately include transfer and novel variants, and treat simulation evidence as one input to calibration — not as proof of real-world readiness.

How it implements the components

Simulation or Case Test fills the manufactured-evidence-and-transfer slice:

  • performance_evidence_set — produces observable, measured performance data from the enacted scenario, safely and on demand.
  • transfer_context_probe — deliberately varies the scenario away from the familiar to test whether competence generalizes rather than being memorized.

It does not compare that evidence to self-assessment or a standard (calibration_gap_map and performance_benchmarkCalibration Exercise and Competency Framework); it manufactures the evidence others calibrate against.

  • Instantiates: Competence Calibration Feedback — manufactures safe, high-fidelity evidence of real performance, including on unfamiliar cases.
  • Sibling mechanisms: Skills Assessment · Supervised Practice · Benchmarked Feedback · Calibration Conversation · Confidence Rating Scale · Calibration Exercise · Exemplar Comparison · Competency Framework · Peer Review · Reflective Error Log · Decision Rights by Competence

References

[1] Line-Oriented Flight Training (LOFT) is a real aviation practice in which crews fly a full, realistic mission scenario in a simulator rather than drilling isolated maneuvers. It is designed to surface how crews actually perform — decision-making, coordination, workload management — under conditions close to real operations.

[2] Transfer of learning refers to whether a skill practiced in one context carries over to another; "near transfer" is to similar situations, "far transfer" to dissimilar ones. Deliberately testing far transfer is how a simulation distinguishes genuine competence from a routine that only works on the practiced case.