Skip to content

Reliability Paradox

Explain why tasks with robust group-level effects (Stroop, IAT) can be useless for ranking individuals: the design minimized within-subjects error for group power without guaranteeing the between-subjects variance that reliability, true-score over total variance, requires.

Core Idea

The reliability paradox, articulated by Hedge, Powell, and Sumner (2018), is the finding that tasks producing robust group-level effects — Stroop, flanker, the IAT — routinely show poor test-retest reliability when scored as individual-differences measures. Reliability equals between-subjects true-score variance over total variance (true-score plus within-subjects error). Paradigms optimized to minimize within-subjects error for group power do not thereby guarantee large between-subjects variance, so the ratio is low even when the group effect is unambiguous. The paradigm is solving the wrong optimization problem.

Scope of Application

The paradox lives within human experimental measurement — the disciplines running repeated-measurement designs with multiple subjects, where paradigms, trials, subjects, and scores make the reliability formula apply unchanged.

  • Cognitive psychology — the originating turf: Stroop, flanker, attentional blink.
  • Implicit social cognition — the IAT, the contested case with weak individual-ranking reliability.
  • Task-based fMRI neuroscience — group-mean maps unstable for ranking individuals.
  • Clinical assessment — patient-vs-control paradigms failing to rank patients on a dimension.
  • Educational and psychometric testing — items discriminating at a cutoff, the home framework.

Clarity

Naming the paradox exposes a category error: treating "this is a well-established effect" as licence to deploy the paradigm for any individual-differences question. It forces "valid for what purpose?" to be answered rather than assumed. The clarity comes from refusing to let two variance components blur into one notion of "a good measure" — the design choices earning group-level power do nothing to guarantee the between-subjects spread correlational work requires, and it sets a computable ceiling: reliability caps any correlation.

Manages Complexity

A scattered record of disappointments — Stroop failing to predict lapses, the IAT failing to predict behaviour — collapses into one structural diagnosis: the paradigm was optimized for one variance partition while the question depends on another. The whole thing runs through the reliability formula, reducing "is this a good measure?" to two variance components and their ratio, with a decidable branch and a precomputable attenuation bound (reliability 0.4 caps the correlation near 0.63) that prunes doomed analyses.

Abstract Reasoning

The paradox licenses a variance-partition attribution dissolving an apparent contradiction into a clean fault localization, a purpose-relative boundary move asking "valid for what?" before deploying any task, a precomputable attenuation ceiling that prunes doomed analyses from one number, and an interventionist move reading remedies (more trials, process-model parameters, a between-subjects-built paradigm) directly off the responsible term of the formula.

Knowledge Transfer

Within human experimental measurement the paradox transfers as mechanism across cognitive psychology, implicit social cognition, fMRI, clinical assessment, and educational testing — the reliability formula, the diagnosis, and the corrective menu hold unchanged because the object is the same repeated-measurement design, and the classical-test-theory vocabulary travels untranslated. Beyond that substrate the insight — within-unit and between-unit variance are different quantities — is carried by the parents variance and measurement_uncertainty. The attenuation ceiling travels literally as the variance apparatus; the paradigm-and-trials machinery stays home.

Relationships to Other Abstractions

Local relationship map for Reliability ParadoxParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Reliability ParadoxDOMAINPrime abstraction: Measurement Uncertainty and Observational Noise — is a decomposition ofMeasurement Unc…PRIME

Current abstraction Reliability Paradox Domain-specific

Parents (1) — more general patterns this builds on

  • Reliability Paradox is a decomposition of Measurement Uncertainty and Observational Noise Prime

    The Reliability Paradox is a measurement-uncertainty failure localized by separating within-unit error from between-unit true-score variance.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Reliability Paradox sits in a sparse region of the domain-specific corpus (71st percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Social Perception & Self-Referential Bias (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12