Pre/Post Capacity Assessment¶
Measurement protocol — instantiates Progressive Stressor Conditioning
Measures capacity before and after a conditioning block — including a delayed transfer test — so real durable gains are separated from momentary performance.
Because productive stress often lowers immediate performance while building durable capacity, you cannot tell whether a conditioning program worked by watching people during it. Pre/Post Capacity Assessment is the measurement instrument that settles the question: it defines the target capacity in testable terms, takes a clean baseline before exposure, and then re-measures after recovery — crucially with a delayed transfer test on novel material, not just a repeat of the trained task. Its defining move is quantifying the delta the right way: comparing a before-reading to a properly delayed, transfer-oriented after-reading, so a genuine capacity gain is distinguished from momentary fluency, fatigue, or the flattering bump of having just practiced the exact test. It reports a measured change; it does not run the debrief that decides what the change means.
Example¶
A surgical residency wants to know whether a new simulation curriculum actually builds laparoscopic skill or just makes residents better at the simulator. It runs a Pre/Post Capacity Assessment. First the target capacity is defined in measurable terms — completion time, error count, and economy of motion on a standardized task, plus successful performance on a task the residents have never drilled. Every resident is measured on this battery before the curriculum begins: the baseline. They then complete the conditioning block. The post-measurement is deliberately taken after a recovery gap, and it includes a transfer task — a procedure structurally different from the drilled one — because the point is durable, generalizable skill, not memorized reps of one exercise. The result is a defensible delta per resident: this cohort improved on the transfer task, that resident's gain was mostly on the drilled task and may not generalize. The assessment does not itself prescribe the training or interpret why a resident stalled; it produces the trustworthy before-and-after numbers those judgments rest on.
How it works¶
- Define capacity as a testable battery. Turn the target capability into concrete, scorable measures — including at least one that generalizes beyond the trained task.
- Take a clean baseline. Measure before exposure, under conditions comparable to the post-test, so the two readings are actually comparable.
- Re-measure after recovery, with transfer. Take the post-reading once fatigue has cleared and include novel-context tasks, so you catch durable capacity rather than momentary or task-specific performance.
- Report the delta with its caveats. State the change and the threats to it (practice effects, who dropped out), rather than a bare "improved."
Tuning parameters¶
- Baseline–post comparability — how tightly matched the two testing conditions are; tight matching makes the delta trustworthy but is costly to standardize, loose matching invites confounds.
- Transfer distance — how far the post-test departs from the trained task; near transfer is easy to show but may just be practice, far transfer is the real prize but harder to demonstrate.
- Post-test delay — how long after exposure the re-measurement is taken; longer delays reveal true durability but risk decay and dropout, immediate re-tests flatter with transient gains.
- Measurement burden — how extensive the battery is; richer batteries catch more but cost time and can themselves fatigue or annoy the people being measured.
When it helps, and when it misleads¶
Its strength is that it is the antidote to the archetype's central illusion — treating immediate performance as durable learning — by insisting the gain be shown on a delayed, transfer-oriented re-measurement rather than on how fluent someone looked mid-training. Its failure mode is attributing improvement to the program that actually came from something else. The sharpest trap is regression to the mean: people selected or tested when they happened to score low will tend to score higher next time regardless of any intervention, so a naive pre/post can manufacture an effect from noise.[1] Practice effects (the post-test is easier because it's the second time), selective dropout, and reactivity are the other usual culprits. The classic misuse is a single pre/post with no control and no delay, presented as proof the training works. The guard is comparability, a real delay, transfer tasks, and honesty about what the delta cannot rule out.
How it implements the components¶
baseline_performance_and_capacity_measure— its signature: the clean, pre-exposure reading that every later gain is measured against.transfer_and_durability_test— the delayed, novel-context re-measurement that distinguishes durable, generalizable capacity from momentary or task-specific performance.target_capacity_definition— it operationalizes the target capability into a concrete, scorable battery so "did capacity rise?" becomes an answerable question.
It quantifies the change but does not interpret or store it: running the facilitated debrief that decides what the change means and writing it into a reusable record — coach_or_feedback_role and gain_retention_record — is After-Action Gain Harvest's. The assessment measures; the harvest makes meaning.
Related¶
- Instantiates: Progressive Stressor Conditioning — it is the measurement backbone that tells whether a conditioning block produced durable capacity.
- Sibling mechanisms: After-Action Gain Harvest · Desirable Difficulty Task Design · Fatigue and Maladaptation Dashboard · Spaced Retrieval and Interleaving Plan · Deload or Recovery Cycle · Graduated Exposure Ladder · Hormetic Microdose Protocol · Consented Challenge Contract · Progressive Overload Protocol
Editorial Notes¶
Form Classification¶
Form family: Assessment, Review & Assurance
Rationale: Pre/Post Capacity Assessment operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it measures capacity before and after a conditioning block — including a delayed transfer test — so real durable gains are separated from momentary performance.
Independent corroboration: The frozen evidence defines Pre/Post Capacity Assessment as 'Measures capacity before and after a conditioning block — including a delayed transfer test — so real durable gains are separated from momentary performance', so its operative form is Assessment, Review & Assurance.
Nearest alternative: Experiment, Test & Rehearsal — Pre/Post Capacity Assessment includes features of an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation, but its defining operation is a bounded evaluation of existing evidence or work that produces a finding or disposition.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Sport Science & Kinesiology
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Specialized
Rationale: Pre/post testing with delayed transfer is rooted in training science and performance conditioning.
Related originating lineages:
- Medicine & Healthcare — Rehabilitation and clinical functional assessment materially shape safe capacity measurement.
- Statistics & Experimental Design — Statistics contributes measurement design and separation of durable change from transient noise.
Review resolution: Both blind reviewers agree that sport science is the primary origin. Reconciliation resolves alternate origin disagreement. Formative alternate lineages are retained as statistics_experimental_design, medicine_healthcare; later breadth of use is recorded separately as domain_reach=specialized, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.
Review outcome: Reconciled after independent review; high confidence.
References¶
[1] Campbell, D. T., & Stanley, J. C. Experimental and Quasi-Experimental Designs for Research. Rand McNally (1963). Campbell and Stanley show that groups selected at low extreme scores tend to score higher later regardless of genuine intervention effects, so naive pre/post comparisons can manufacture an apparent effect. They do not call regression to the mean the single sharpest trap. registry ↩