Dynamic Source-Reliability Scorecard¶
Governance scorecard — instantiates Adaptive Precision-Weighted Signal Fusion
A maintained, human-readable rating of each source's reliability across multiple axes, updated as sources perform, that governs how much mixed or qualitative evidence should count.
A Dynamic Source-Reliability Scorecard is the archetype's home for qualitative and mixed evidence, where no clean variance exists to invert. It is a maintained artifact — a register of every source rated on several distinct reliability axes — that people can read, contest, and act on. The idea that makes it this mechanism is that reliability here is a governed, multi-axis, human-legible rating rather than a computed number: a source is not "0.7 reliable" but, say, "usually reliable / independent origin / stale on this topic," and those tiers are revised as the source's track record and situation change. Its distinctive output is not a fused estimate but a label a stakeholder can trust or challenge — the connective tissue that lets an analyst decide how much a given report should count without pretending the judgment is exact.
Example¶
An intelligence desk is weighing several reports on the location of a shipment: a long-trusted human source, a satellite image, an intercepted message, and an anonymous tip. The desk keeps a scorecard rating each source on separate axes — historical reliability, independence of origin, and current-context fit — using tiers rather than false decimals, in the spirit of the long-standing military convention that scores a source's reliability on one scale and each report's credibility on another.[n1] The trusted human source is rated highly on reliability but flagged stale on this region, where they have not operated in a year; the anonymous tip is rated low but independent. The scorecard is dynamic: when the human source's last three regional reports proved wrong, an analyst downgrades their context-fit tier and records why. Downstream, these labels tell everyone how much each report should weigh — and the anonymous tip, though low-rated, is preserved because it is the only independent voice, rather than being dismissed for lacking pedigree.
How it works¶
- Register every source. Maintain an inventory of sources with provenance and what each is positioned to know.
- Rate on multiple axes. Score reliability, independence, timeliness, context-fit, and known failure modes separately, so a stable-but-biased source and a noisy-but-independent one are distinguishable.
- Revise on evidence. When a source's track record, situation, or behaviour changes, an analyst updates the relevant tier and records the reason, keeping the rating current.
- Publish legible labels. Emit each source's rating as a stakeholder-readable tier and rationale that governs how much its evidence counts.
Tuning parameters¶
- Axis set — which distinct reliability dimensions are tracked. More axes capture nuance but raise the effort to keep every cell current.
- Tier granularity — how many rating levels each axis has. Coarse tiers resist false precision; fine ones risk implying more exactness than the judgment supports.
- Update trigger — what events force a re-rating (a missed call, a context shift, elapsed time). Sensitive triggers keep ratings fresh but demand constant analyst attention.
- Rationale requirement — whether every rating change must record a reason. Mandatory rationale makes the scorecard auditable at the cost of friction.
When it helps, and when it misleads¶
Its strength is legitimacy where math cannot reach: it makes reliability contestable and transparent, keeps multiple failure dimensions from collapsing into one vague "trust" score, and gives non-technical stakeholders a rating they can actually reason about. It is the archetype's answer to evidence that is qualitative, political, or sparse.
It misleads through the pathologies of human rating. A halo effect lets a source's reputation for reliability inflate its credibility on a topic it knows nothing about; tiers left un-revised go stale and keep endorsing a source whose context has moved; and if the axes are not kept distinct, the scorecard quietly becomes a single popularity score dressed up as many.[n1] The guarding discipline is to force periodic re-rating against outcomes, to keep the reliability and context-fit axes rigorously separate, and to treat an unchanged rating on a fast-moving source as a warning rather than a reassurance.
How it implements the components¶
candidate_signal_inventory— the register of sources and their provenance is the scorecard's backbone.signal_quality_profile— the multi-axis rating is a qualitative quality profile, keeping bias, independence, timeliness, and context-fit distinct.context_sensitive_weight_update_trigger— defined events (missed calls, regime shifts, elapsed time) force analysts to re-rate, so reliability does not freeze.stakeholder_interpretation_label— publishes each rating as a human-readable tier and rationale others can trust or challenge.
It rates sources but computes no fused number and runs no decay arithmetic: the numeric precision_or_reliability_weight_rule that ages weights on a schedule is Weight Decay and Refresh Schedule's, and the fused output (fused_estimate_with_uncertainty_state) belongs to the estimators. Its updates are analyst judgment, not the formal held-out feedback_calibration_loop of Cross-Validation Weight Calibration.
Related¶
- Instantiates: Adaptive Precision-Weighted Signal Fusion — the qualitative, governance-side reliability layer.
- Sibling mechanisms: Inverse-Variance Weighting · Bayesian Cue Integration Model · Kalman Filter Update · Weighted Ensemble Estimator · Confidence-Weighted Vote · Sensor-Fusion Pipeline · Weight Decay and Refresh Schedule · Cross-Validation Weight Calibration
Editorial Notes¶
Form Classification¶
Form family: Monitoring, Sensing & Alerting
Rationale: The maintained scorecard repeatedly observes source performance and context, updates multidimensional reliability indicators, and emits current labels that govern evidentiary weight.
Nearest alternative: Record, Log & Register — Revision reasons provide provenance, but the mechanism's primary value is the current, repeatedly refreshed reliability signal rather than an accumulating history of changes.
Review outcome: Adjudicated after independent review; medium confidence.
Origin Attribution¶
Primary origin: Security Studies & Intelligence Analysis
Origin pattern: Single lineage
Present-day reach: Multi-domain
Rationale: Intelligence analysis cohered source-evaluation matrices such as the Admiralty Code, separating source reliability from report credibility and updating both with performance.
Related originating lineages:
- Statistics & Experimental Design — Evidence weighting and calibration supplied quantitative checks on source performance over time.
Review resolution: Security and intelligence is primary because source evaluation is a standing analytic tradecraft; statistics supplies calibration and evidence-weighting discipline.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
[n1] The NATO "Admiralty Code" (or source-evaluation matrix) rates a source's reliability on one lettered scale (A–F) and each report's credibility on a separate numbered scale (1–6), deliberately keeping the two axes independent so a reliable source's dubious report is not automatically believed. It is the canonical model for a multi-axis, human-legible reliability scorecard. ↩a ↩b