Mixed Diagnostic Quiz¶
Assessment — instantiates Interleaved Discrimination Practice
A delayed, unannounced-category assessment that requires classification or method selection in mixed order, to measure whether discrimination survives outside the block.
A Mixed Diagnostic Quiz is an assessment, not a practice loop: its job is to measure discrimination, not to build it. Items from several already-taught target categories are presented in mixed order, after a delay, with no announcement of which category each belongs to, and the test-taker must classify or select a method before any label is given. Because nothing pre-announces the category and the items arrive well after the teaching block, the quiz strips away exactly the contextual crutches that make blocked performance look strong — and reveals whether real discrimination is underneath. It is deliberately not a feedback-rich teaching artifact; it is the honest ruler. Its defining feature is that it functions as a delayed transfer probe: the score it produces answers "can they still tell these apart, cold, and in unfamiliar dress?" rather than "did they follow today's lesson?"
Example¶
A geology instructor's students crush the mineral-identification worksheet at the end of the silicates unit — but the worksheet came right after the silicates lecture, and every specimen was a silicate, so "it's a silicate, now which one" was half-answered before they picked up the streak plate. Three weeks later she runs a Mixed Diagnostic Quiz as a lab practical: twenty numbered stations, each a hand specimen drawn without announcement from silicates, carbonates, sulfides, and oxides taught across the whole term, in scrambled order. At each station a student must, before consulting anything, commit in writing to the mineral and the diagnostic property that decided it — hardness, cleavage, streak, reaction to acid. Several stations are transfer items: a weathered or unusually-colored specimen unlike any seen in class. No feedback is given during the practical; the grade records where discrimination held and where it collapsed — and it turns out half the class had been leaning on the tell that "if it's on the silicates worksheet, it's a silicate." The practical is what made that dependence visible.
How it works¶
- Draw across all taught categories, unannounced. The item set spans the full confusable target space, and no heading, section, or ordering signals a category — the test-taker supplies it.
- Delay from instruction. The quiz is administered well after the relevant teaching, so it probes durable discrimination rather than fresh recall.
- Require the commit before any label. Each item demands a classification or method choice and the deciding property, up front, so the graded act is the discrimination itself.
- Embed transfer items. Some items appear in novel surface dress — unfamiliar context, atypical presentation — to probe whether the distinction generalizes beyond trained exemplars.
- Score, do not coach. The output is a measurement of where discrimination held and failed, routed to the designer to revise practice — not returned as in-the-moment teaching.
Tuning parameters¶
- Delay length — time between instruction and quiz. Longer delays give a more honest read on durability but can conflate discrimination failure with plain forgetting.
- Transfer-item fraction — share of items in novel dress. More transfer items test generalization harder but raise difficulty and grading nuance.
- Distractor calibration — how confusable the wrong options or nearby categories are. Tight distractors give a sensitive measure; loose ones let shallow cues pass.
- Scoring scheme — whether partial credit is given for the right method with a wrong final answer, or discrimination is scored separately from execution. Separating them isolates which competence failed.
- Format distance from practice — how much the quiz's surface differs from the practiced format. Greater distance guards against measuring format familiarity; too much distance measures unfamiliarity itself.
When it helps, and when it misleads¶
Its strength is that it is the one instrument that catches the archetype's core illusion — block-inflated fluency — before the field does, by scoring discrimination under the unannounced, delayed, mixed conditions that real use imposes. Its logic is transfer-appropriate processing: performance is best predicted by the match between how knowledge was practiced and how it is tested, so a quiz whose conditions mirror real use is the valid measure, and a blocked quiz is not.[n1]
Its failure mode is subtle: if the quiz format hugs the practice format too closely, it measures format familiarity rather than transfer, and a class can ace it while still failing in the field. Run repeatedly and high-stakes, it also decays into something to teach to — practice narrows to quiz-shaped items and the measure stops being independent of what it measures. The classic misuse is quietly reintroducing the block: grouping the quiz by category or announcing the topic, which re-supplies the crutch and inflates the score into meaninglessness. The guarding discipline is to keep categories unannounced and interleaved, rotate transfer items so the format cannot be memorized, and keep the quiz separate from the practice stream so it stays an honest external check.
How it implements the components¶
discrimination_target_set— the item set spans the full confusable target space and presents it unannounced, forcing the test-taker to supply the category rather than receive it.retrieval_or_performance_prompt— the required pre-label commit to a category and its deciding property makes the discrimination act the thing being measured.delayed_retention_and_transfer_probe— administered after a delay with novel-dress transfer items, the quiz is the archetype's honesty check on durable, generalizable discrimination.
It does NOT implement the comparison_feedback_loop — a diagnostic quiz measures and withholds in-the-moment cue teaching; the feedback that turns a miss into boundary learning is Mixed Problem Set's (which shares the target-set and prompt but teaches rather than measures), and cue-level correction on confusable pairs is Near-Miss Case Rotation's. Sequencing and adaptation are Adaptive Interleaving Scheduler's.
Related¶
- Instantiates: Interleaved Discrimination Practice — supplies the delayed, mixed measurement that keeps the whole design honest.
- Consumes: Near-Miss Case Rotation can supply confusable distractors for sensitive items.
- Sibling mechanisms: Mixed Problem Set · Shuffled Practice Deck · Near-Miss Case Rotation · Interleaved Repertoire Practice · Adaptive Interleaving Scheduler · Contrastive Example Sequence · Alternating Context Drill
Editorial Notes¶
Form Classification¶
Form family: Assessment, Review & Assurance
Rationale: The delayed mixed quiz evaluates durable discrimination across confusable categories and returns evidence of whether the learner can classify without advance labels.
Nearest alternative: Experiment, Test & Rehearsal — Items include novel transfer cases, but the defining output is a finding about existing mastery rather than learning through deliberately varied practice.
Review outcome: Adjudicated after independent review; high confidence.
Origin Attribution¶
Primary origin: Education & Pedagogy
Origin pattern: Convergent development
Present-day reach: Multi-domain
Rationale: Mixed, delayed assessment is an established instructional-design technique for testing discrimination and transfer beyond blocked practice.
Related originating lineages:
- Psychology — Cognitive psychology supplied transfer-appropriate processing and memory evidence explaining why mixed retrieval is diagnostically stronger.
Review resolution: Both independent reviews agree on primary origin education_pedagogy; reconciliation resolves secondary fields (origin_mode_disagreement, domain_reach_disagreement). Alternate origins retained (psychology) are the union of reviewer-supported formative lineages with explicit rationales, not a list of later application domains. Present-day breadth is represented separately as domain_reach=multi_domain; origin_mode=convergent records the historical relationship among lineages. Confidence is conservatively reconciled to high, and encyclopedia_synthesis=false preserves either reviewer's finding that the encyclopedia generalized the mechanism.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
Keep this mechanism out of the practice loop it evaluates. The moment the same mixed quiz is used both to train and to certify, its independence is gone and its score stops meaning "discrimination survived" and starts meaning "the learner has seen this test." Pair it with, but never merge it into, the practice mechanisms.
[n1] Transfer-appropriate processing (Morris, Bransford, and Franks) holds that retention is governed by the overlap between the cognitive processing done at study and at test — recall is best when test conditions reinstate the kind of processing practice demanded. It is the reason a mixed, unannounced quiz validly measures discrimination while a blocked quiz measures the block. ↩