Summary-Only Reader Test¶
Method — instantiates Summary-Substance Alignment Audit
Puts the summary in front of readers who never see the body and measures what they conclude, catching the gap between what the summary says and what a summary-only audience takes away.
A Summary-Only Reader Test is empirical. It puts the summary in front of people who never see the body and measures what they actually conclude, then compares those conclusions to what the substance supports. Where every other mechanism inspects the text, this one inspects the reader: it treats the summary's real meaning as whatever a representative summary-only audience takes away, and catches the gap between what the summary technically says and what it causes people to believe. A summary can pass every text audit — accurate, fully qualified, cleanly traced — and still fail here, because the failure lives in the inference, not the words.
Example¶
A lender's product team writes a one-page "Key Facts" summary for a variable-rate mortgage. It is accurate line by line. The Summary-Only Reader Test recruits a panel resembling actual applicants, gives them only the Key Facts sheet — never the full contract — and asks what they would expect: What is your payment in year three if rates rise two points? Can the rate change in the first year? A majority answer that the rate is fixed for two years, because the sheet's phrase "introductory rate period" reads, to a non-specialist, as a rate guarantee — though the contract only fixes the margin, not the rate. No wording was false; the inference was. The test surfaces the misread that a text audit cannot see, and the sheet is rewritten around the belief it actually produces rather than the one it technically permits.
How it works¶
The method measures inference, not alignment: (1) define the summary-only audience — who really acts on the summary alone, and with what background; (2) show a matched panel the summary in isolation; (3) elicit the beliefs and intended actions it produces, through recognition, recall, or "what would you do" questions; (4) compare those against what the substance supports and flag the material misreads. The unit of failure is a wrong reader inference drawn from true words — the distortion no comparison of text to text can find, because both texts are fine.
Tuning parameters¶
- Panel representativeness — how closely testers match the real audience. A panel of insiders will not misread the way outsiders do, hiding the very gap you are testing for.
- Question design — recognition ("is the rate fixed?"), free recall ("what did this say?"), or intended action ("what would you do?"). Each catches a different layer of misunderstanding.
- Isolation strictness — summary only, or summary-first-then-body; strict isolation reproduces the real summary-only reader, the person who never clicks through.
- Misread threshold — what share of readers drawing the wrong inference counts as a defect worth a rewrite.
- Iteration — one-shot, or test–rewrite–retest until the misread rate falls below the threshold.
When it helps, and when it misleads¶
Its strength is catching the distortion that lives in the reader rather than the text — the technically-true summary that predictably misleads — which no word-to-word comparison can find. Its failure mode is that results are only as good as the panel's resemblance to the real audience; a convenient panel of colleagues will sail through a sheet that baffles actual customers. The classic misuse is running it as theater — to certify a summary already shipped, or leading the questions until readers give the "right" answers. The discipline is a representative panel, neutral questions, and a pre-set misread threshold, applied before publication. The bias it externalizes is the curse of knowledge: authors who know the substance systematically misjudge what a naïve reader will infer from the summary.[1]
How it implements the components¶
Summary-Only Reader Test fills the reader-facing, empirical side of the archetype:
summary_only_audience_profile— defines who actually acts on the summary alone, and with what expertise; the panel is drawn to match it.reader_inference_test_panel— the panel and protocol that elicit and measure real readers' conclusions from the summary in isolation.
It measures what readers infer; it does not locate where in the body a claim is (or is not) supported (Summary-Claim Traceability Matrix) or inventory the qualifiers the summary dropped (Qualifier-Drop Scan) — those are text audits; this is a reader audit.
Related¶
- Instantiates: Summary-Substance Alignment Audit — supplies the empirical, audience-side check the text audits cannot.
- Sibling mechanisms: Summary-Claim Traceability Matrix · Qualifier-Drop Scan · Press-Release Claim Review · Certainty & Causality Inflation Check · Quote-Snippet Context Window · Executive-Summary Caveat Budget
Notes¶
This is the only mechanism in the set that can catch a distortion where every word is true and every qualifier is present, yet readers still leave with the wrong belief. A summary that is perfectly clean by every text audit can still fail the reader test — which is why it is worth running last, on summaries the other mechanisms have already passed, to catch what alignment-checking structurally cannot.
References¶
[1] The curse of knowledge — the well-documented bias whereby someone who knows a thing struggles to model what a person without that knowledge will understand. It is precisely why an author's own read of a summary is a poor proxy for a summary-only reader's, and why the inference must be measured rather than assumed. ↩