Known-Groups or Contrast-Case Test¶
A method — instantiates Construct–Proxy–Signal Validity Alignment
Checks that the measure separates groups already known to differ on the construct — and that the separation isn't explained by a confound the groups also differ on.
If a measure captures a construct, it should behave the way the construct predicts when you already know the answer. Known-Groups or Contrast-Case Test exploits exactly that: it picks groups (or paradigm cases) that theory says must differ on the construct and checks that the measure detects the difference in the right direction and size. Its defining and easily-skipped second half is ruling out the alternative explanation — because groups that differ on the construct almost always differ on other things too, the test only counts if the separation survives when a plausible confound is matched away. It is the external, criterion-facing counterpart to the internal checks: not "do the items hang together?" but "does the score move when the world moves?" — and "for the right reason?"
Example¶
A brief depression screener is meant to be handed to primary-care patients. The known-groups test compares a group with clinician-diagnosed major depressive disorder against a matched community sample: the screener should score markedly higher in the diagnosed group, and it does. But the sharper test is the confound check. The team also administers it to a group that is severely sleep-deprived and physically exhausted but not depressed — and finds the screener scores them high too, because its somatic items (fatigue, poor sleep, low energy) fire on exhaustion. The group contrast that "worked" was partly tracking a surrogate. The screener discriminates depression from the general population, but not depression from fatigue — a boundary the plain known-groups comparison would have hidden.
How it works¶
- Choose groups the construct must separate. Identify populations or cases theory says differ sharply on the construct, and predict the direction and rough magnitude before testing.
- Match away the confounds. Design or select the contrast so the groups differ on the construct while being comparable on plausible nuisance variables — otherwise a surrogate can produce the separation.
- Test the boundary case. Add a group that differs on a likely confound but not the construct; if the measure still separates it, the signal is surrogate-driven, not construct-driven.
Tuning parameters¶
- Contrast extremity — how far apart the known groups are; extreme groups make any measure look valid, graded contrasts are a harder, fairer test.
- Confound matching tightness — how carefully the groups are equated on nuisance variables; looser matching inflates apparent validity.
- Predicted effect size — whether success is any difference or a pre-specified magnitude; a stated magnitude is a stronger test than mere direction.
- Boundary-group inclusion — whether a confound-only contrast group is added; without it the surrogate risk goes unmeasured.
When it helps, and when it misleads¶
Its strength is that it is intuitive and, when the contrast is clean, powerful: a measure that cannot tell apart groups everyone agrees differ has a serious problem, full stop. It ties the score to the external world rather than to its own internals.
Its central weakness is confounding — real groups differ on many dimensions at once, so an unmatched contrast can "validate" a measure that is tracking a correlate (age, education, acute distress) rather than the construct. Extreme-group designs make this worse by flattering weak measures. The classic misuse is choosing maximally different groups so that separation is guaranteed and reported as validity. The discipline is to match confounds, prefer graded over extreme contrasts, and always include the boundary case that could expose a surrogate.[n1]
How it implements the components¶
convergent_discriminant_check— the group separation is criterion evidence that the score converges with real, independently-known construct differences.confound_and_surrogate_boundary— the matched-confound and boundary-case design is precisely what draws the line between construct-driven and surrogate-driven separation.
It does not analyze internal item structure or method variance (that is Factor-Structure or Latent-Model Check and the Multi-Trait Multi-Method Matrix), nor does it scope the claim or state its limits (that is the Construct Validity Argument and Validity Limitation Memo).
Related¶
- Instantiates: Construct–Proxy–Signal Validity Alignment — it supplies external, criterion-facing evidence that the measure tracks real construct differences.
- Sibling mechanisms: Multi-Trait Multi-Method Matrix · Factor-Structure or Latent-Model Check · Construct Validity Argument · Content-Domain Review Panel · Construct-to-Proxy Traceability Table · Cognitive Interview or Response-Process Probe · Measurement Invariance Audit · Proxy Drift and Goodhart Audit · Validity Limitation Memo
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Known-Groups or Contrast-Case Test operates as a bounded trial, probe, simulation, or rehearsal that generates evidence from performance because it checks that the measure separates groups already known to differ on the construct — and that the separation isn't explained by a confound the groups also differ on
Independent corroboration: The frozen evidence defines Known-Groups or Contrast-Case Test as 'Checks that the measure separates groups already known to differ on the construct — and that the separation isn't explained by a confound the groups also differ on', so its operative form is Experiment, Test & Rehearsal.
Review outcome: Independent reviewer agreement; high confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Multi-domain
Rationale: Psychometrics and statistical validation developed known-groups tests for construct validity and confound-resistant discrimination.
Related originating lineages:
- Psychology — Measurement theory in psychology supplied the construct-validity problem and paradigm-group designs.
Review outcome: Independent reviewer agreement; high confidence.
Notes¶
[n1] Known-groups validity — showing a measure distinguishes groups already known to differ on the construct — is a standard technique; its value stands or falls on whether the groups differ only on the construct, which is why matched confounds and boundary cases matter. ↩