Skip to content

Blind Pattern Comparison Round

Blinded comparison — instantiates Independent Convergence Evidence Appraisal

Has independent judges rate whether candidate solution-shapes truly match, with lineage identities and the convergence claim masked, so the similarity score is not manufactured by expectation.

Blind Pattern Comparison Round attacks a bias the other mechanisms leave open: the observer's own eagerness to see convergence. Its defining move is masking the raters. Several judges are given the solution-shapes stripped of labels — no lineage names, no dates, no "these are supposed to match" framing — and asked independently to say which shapes are genuinely alike and how alike. Only afterward are the masks lifted and the ratings compared to the convergence hypothesis. The score that comes out is a de-biased measure of how strong the shape-match really is, because it was formed without knowing which answer would flatter the story. The round also watches the raters themselves: if masked, independent judges agree too perfectly, that near-unanimity is itself suspicious — a sign the shapes were obvious duplicates, the raters weren't independent, or the comparison was too easy to be informative.

Example

A linguist is assessing the claim that unrelated languages "independently converge" on subject-verb-object word order because SVO fits a real processing pressure. The worry is that anyone who already believes the convergence will read fuzzy structures as matching. So she runs a blind round. Coders receive sentence structures rendered in a neutral notation, with language names, families, and the SVO hypothesis removed, and are asked simply to cluster the structures by similarity and rate each cluster's tightness.

When the masks come off, two things emerge. First, some structures the proponent had counted as "SVO matches" get clustered apart by the blind coders, because the masked comparison did not let expectation paper over real differences — the de-biased match is weaker than claimed. Second, on a subset the coders agree almost perfectly, and that perfection is a flag, not a triumph: those cases turn out to share a documented contact history, so their "independent" agreement was dependency in disguise. The round hands back a lowered, cleaner strength score and a false-unanimity note on the suspiciously tidy subset.

How it works

  • Strip the anchors. Remove lineage identity, timing, source, and the convergence hypothesis from the material the raters see, leaving only the shapes.
  • Rate independently, then reconcile. Multiple judges score similarity without conferring; only afterward are scores compared, so agreement is earned rather than coordinated.
  • Unmask and map to the hypothesis. Reveal identities and check whether the blind similarity ratings actually line up with the claimed convergence.
  • Read the agreement pattern for danger. Treat near-perfect blind agreement as a prompt to check for dependency or triviality, not as extra confirmation.

Tuning parameters

  • Masking depth — how much identifying context is stripped. Deeper masking removes more bias but can make shapes hard to interpret; too little leaks the answer.
  • Number and independence of raters — more genuinely separate judges give a more trustworthy score but cost coordination and are harder to keep from conferring.
  • Similarity scale — binary match/no-match versus a graded resemblance score. Grades preserve nuance; binaries force cleaner but blunter agreement statistics.
  • Unanimity-alarm threshold — how tight rater agreement must be before it is treated as suspicious rather than reassuring. Set it loose and dependency hides; set it tight and ordinary agreement gets over-flagged.

When it helps, and when it misleads

Its strength is that it removes the one contaminant the paper trail cannot — the appraiser's expectation — and it is the only mechanism here that treats too much agreement as a warning. That inversion is the point: in genuinely noisy independent observation, perfect unanimity is less likely than modest disagreement, so flawless agreement often signals a hidden common cause rather than a strong signal.[1]

Its failure mode is that masking can destroy the very information that makes a match meaningful: strip too much context and raters compare shapes shorn of the functional detail that distinguishes deep structural similarity from surface resemblance, producing a clean but hollow score. The classic misuse is running a single rater "blind" and calling it objective, when one masked judge is just one biased judge with a blindfold. The guarding discipline is to use several independent raters, mask only identity and hypothesis (not function), and treat the round as one input to strength rather than the last word.

How it implements the components

  • convergence_strength_assessment — the round's output is a de-biased strength score: how much the shape-match should move belief, measured without the raters knowing which answer helps the story.
  • false_unanimity_warning — by watching for suspiciously perfect blind agreement, it raises the flag that near-unanimity may signal dependency or triviality rather than strength.

It does NOT tabulate the shapes, pressures, and independence evidence into the working grid (solution_shape_abstraction, shared_pressure_profile, lineage_independence_map) — that is its nearest twin, Convergence Evidence Matrix; the matrix organizes the evidence, while this round supplies the masked judgment of how strong the shape-match within it actually is.

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Independent judges rate masked candidate shapes, reconcile their scores, then unmask lineage to determine whether claimed convergence is supported, so its operative form is blinded assessment.

Nearest alternative: Experiment, Test & Rehearsal — Masking removes bias but does not perturb the candidate patterns; judges evaluate fixed evidence rather than generate outcomes through a trial.

Review outcome: Adjudicated after independent review; high confidence.

Origin Attribution

Primary origin: Statistics & Experimental Design

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Independent masked judging of hypothesized similarities is an experimental-design safeguard against confirmation and expectancy bias.

Related originating lineages:

Review resolution: Statistics is the agreed primary lineage through independent masked comparison before a hypothesis can anchor raters. Cognitive pattern recognition and qualitative comparative practice shape what is judged; the protocol has multi-domain reach despite its specialized review setting.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; high confidence.

References

[1] Gunn, L. J., Chapeau-Blondeau, F., McDonnell, M. D., Davis, B. R., Allison, A., & Abbott, D. "Too Good to Be True: When Overwhelming Evidence Fails to Convince". Proceedings of the Royal Society A 472(2187), 20150748 (2016). Shows probabilistically that, when observations are subject to even rare systemic failure, sufficiently unanimous results can be less credible than modest disagreement and can indicate a hidden failure state. registry