Skip to content

Evidence Weighting Rubric

Scoring tool — instantiates Salience-Significance Decoupling

Scores evidence against explicit significance criteria fixed before the evidence is seen, so vividness cannot smuggle in weight it has not earned.

Left to itself, evidence gets weighted by how strongly it lands — the vivid testimony outweighs the dull dataset. Evidence Weighting Rubric overrides that by fixing what counts as weighty before any particular item is in view, then scoring each piece of evidence against those criteria — diagnosticity, source reliability, corroboration, base-rate consistency, relevance to the actual question — rather than against how much it grabs. Its defining feature is the pre-committed criteria and the gate that follows: because the significance standard is set in advance and the decision is required to move on the weighted score, a striking anecdote cannot tip the conclusion by force of vividness alone. Where other mechanisms find missing cases or restore baselines, this one supplies the ruler — the explicit account of what makes evidence matter — and enforces its use.

Example

An intelligence desk is assessing whether a supplier disruption is imminent. Reports stream in, and one is unforgettable: a dramatic first-hand account, richly detailed, already recirculated widely. Its vividness is doing quiet work on everyone who reads it. The rubric was set before the reports arrived: each item is scored on diagnosticity (does it distinguish this hypothesis from its rivals?), source reliability, independent corroboration, and consistency with the base rate for such disruptions.

Scored that way, the vivid account rates low — it is consistent with almost any hypothesis, so it barely discriminates, however gripping it reads. A dry logistics dataset that few noticed rates high, because it separates the live hypothesis from the alternatives. The gate then binds: the desk's assessment moves on the weighted total, so the memorable anecdote cannot carry the call on vividness. The output is a picture in which attention and weight have finally been pried apart — what got noticed and what counts are two different columns.

How it works

  • Fix the criteria first. Write down what makes evidence significant for this question — diagnosticity, reliability, corroboration, relevance — before the evidence is in hand, so the standard cannot be bent to fit a favourite item.
  • Score against the register, not the impression. Rate each piece on the criteria, explicitly setting aside how vivid, recent, or emotionally loud it is; salience is not one of the columns.
  • Gate the decision on weight. Require the conclusion to move on the aggregated weight, so a high-salience, low-weight item is structurally barred from driving action before it clears the bar.
  • Show the scores. Keep the per-item scoring visible, so a challenge lands on a specific criterion ("this isn't as diagnostic as scored") rather than on a gut sense.

Tuning parameters

  • Criteria set and weights — which dimensions count and how much each is worth. The highest-leverage dial and the one to fix first: reweighting after seeing the evidence is how a rubric gets quietly gamed.
  • Diagnosticity emphasis — how much the rubric rewards evidence that discriminates between hypotheses versus merely fits one. Emphasizing it is the strongest guard against vivid-but-uninformative material.
  • Gate strictness — how firmly the decision is bound to the weighted score. A hard gate resists salience but can feel mechanical; a soft one restores judgment and reopens the door to vividness.
  • Scoring grain — a quick high/medium/low pass versus a fine numeric scale. Finer scoring discriminates better but invites false precision and slows the work.

When it helps, and when it misleads

Its strength is that it separates noticed from weighed by construction — the pre-committed criteria and the gate make it structurally hard for salience to masquerade as significance, and the visible scores turn arguments about a hunch into arguments about a named criterion.[1] It is at its best where vivid evidence routinely crowds out the informative-but-dull: analysis, review, triage.

Its failure modes are the usual ones for a scoring instrument. A rubric with the wrong criteria confidently weights the wrong things, its explicitness lending false authority to a bad standard. It can be run backwards — criteria or weights tuned after the fact to justify the conclusion already reached — which is precisely what pre-committing them is meant to prevent, and precisely what erodes when nobody checks. And it tends to under-weight evidence that resists scoring, quietly zeroing out what the criteria failed to anticipate. The discipline is to fix the criteria before the evidence, keep the scores visible and contestable, and treat the weighted total as a structured argument whose criteria must hold — not a verdict that ends the conversation.

How it implements the components

This rubric realizes the define-and-enforce-significance side of the archetype — the ruler and the gate, not the search for cases:

  • significance_criteria_register — it is this register: the explicit, pre-committed account of what makes evidence weighty for the question at hand.
  • decision_weight_gate — it binds the decision to the weighted score, so prominence cannot drive action until an item has cleared the significance bar on its merits.

It sets and applies the weighting criteria but does not gather the base rates it weights — those come from Base-Rate Visibility Panel — nor surface the disconfirming cases to be weighed, which Counterexample Surface Scan supplies.

Notes

The rubric is the significance-defining core several other mechanisms lean on: Dashboard Salience Calibration allocates display prominence against the criteria this register fixes, and Counterexample Surface Scan needs it to decide which surfaced counter-cases are worth weighing. Keeping the criteria explicit and pre-committed is what lets those consumers inherit a standard instead of improvising one per decision.

References

[1] Analysis of Competing Hypotheses (Richards Heuer) — a structured analytic technique that scores each piece of evidence by how well it discriminates among rival hypotheses rather than by how strongly it supports a favoured one. It is the canonical instance of weighting by diagnosticity, which is why diagnosticity leads the criteria above.