Skip to content

Dynamic Source-Reliability Scorecard

Governance scorecard — instantiates Adaptive Precision-Weighted Signal Fusion

A maintained, human-readable rating of each source's reliability across multiple axes, updated as sources perform, that governs how much mixed or qualitative evidence should count.

A Dynamic Source-Reliability Scorecard is the archetype's home for qualitative and mixed evidence, where no clean variance exists to invert. It is a maintained artifact — a register of every source rated on several distinct reliability axes — that people can read, contest, and act on. The idea that makes it this mechanism is that reliability here is a governed, multi-axis, human-legible rating rather than a computed number: a source is not "0.7 reliable" but, say, "usually reliable / independent origin / stale on this topic," and those tiers are revised as the source's track record and situation change. Its distinctive output is not a fused estimate but a label a stakeholder can trust or challenge — the connective tissue that lets an analyst decide how much a given report should count without pretending the judgment is exact.

Example

An intelligence desk is weighing several reports on the location of a shipment: a long-trusted human source, a satellite image, an intercepted message, and an anonymous tip. The desk keeps a scorecard rating each source on separate axes — historical reliability, independence of origin, and current-context fit — using tiers rather than false decimals, in the spirit of the long-standing military convention that scores a source's reliability on one scale and each report's credibility on another.[n1] The trusted human source is rated highly on reliability but flagged stale on this region, where they have not operated in a year; the anonymous tip is rated low but independent. The scorecard is dynamic: when the human source's last three regional reports proved wrong, an analyst downgrades their context-fit tier and records why. Downstream, these labels tell everyone how much each report should weigh — and the anonymous tip, though low-rated, is preserved because it is the only independent voice, rather than being dismissed for lacking pedigree.

How it works

  • Register every source. Maintain an inventory of sources with provenance and what each is positioned to know.
  • Rate on multiple axes. Score reliability, independence, timeliness, context-fit, and known failure modes separately, so a stable-but-biased source and a noisy-but-independent one are distinguishable.
  • Revise on evidence. When a source's track record, situation, or behaviour changes, an analyst updates the relevant tier and records the reason, keeping the rating current.
  • Publish legible labels. Emit each source's rating as a stakeholder-readable tier and rationale that governs how much its evidence counts.

Tuning parameters

  • Axis set — which distinct reliability dimensions are tracked. More axes capture nuance but raise the effort to keep every cell current.
  • Tier granularity — how many rating levels each axis has. Coarse tiers resist false precision; fine ones risk implying more exactness than the judgment supports.
  • Update trigger — what events force a re-rating (a missed call, a context shift, elapsed time). Sensitive triggers keep ratings fresh but demand constant analyst attention.
  • Rationale requirement — whether every rating change must record a reason. Mandatory rationale makes the scorecard auditable at the cost of friction.

When it helps, and when it misleads

Its strength is legitimacy where math cannot reach: it makes reliability contestable and transparent, keeps multiple failure dimensions from collapsing into one vague "trust" score, and gives non-technical stakeholders a rating they can actually reason about. It is the archetype's answer to evidence that is qualitative, political, or sparse.

It misleads through the pathologies of human rating. A halo effect lets a source's reputation for reliability inflate its credibility on a topic it knows nothing about; tiers left un-revised go stale and keep endorsing a source whose context has moved; and if the axes are not kept distinct, the scorecard quietly becomes a single popularity score dressed up as many.[n1] The guarding discipline is to force periodic re-rating against outcomes, to keep the reliability and context-fit axes rigorously separate, and to treat an unchanged rating on a fast-moving source as a warning rather than a reassurance.

How it implements the components

  • candidate_signal_inventory — the register of sources and their provenance is the scorecard's backbone.
  • signal_quality_profile — the multi-axis rating is a qualitative quality profile, keeping bias, independence, timeliness, and context-fit distinct.
  • context_sensitive_weight_update_trigger — defined events (missed calls, regime shifts, elapsed time) force analysts to re-rate, so reliability does not freeze.
  • stakeholder_interpretation_label — publishes each rating as a human-readable tier and rationale others can trust or challenge.

It rates sources but computes no fused number and runs no decay arithmetic: the numeric precision_or_reliability_weight_rule that ages weights on a schedule is Weight Decay and Refresh Schedule's, and the fused output (fused_estimate_with_uncertainty_state) belongs to the estimators. Its updates are analyst judgment, not the formal held-out feedback_calibration_loop of Cross-Validation Weight Calibration.

Editorial Notes

Form Classification

Form family: Monitoring, Sensing & Alerting

Rationale: The maintained scorecard repeatedly observes source performance and context, updates multidimensional reliability indicators, and emits current labels that govern evidentiary weight.

Nearest alternative: Record, Log & Register — Revision reasons provide provenance, but the mechanism's primary value is the current, repeatedly refreshed reliability signal rather than an accumulating history of changes.

Review outcome: Adjudicated after independent review; medium confidence.

Origin Attribution

Primary origin: Security Studies & Intelligence Analysis

Origin pattern: Single lineage

Present-day reach: Multi-domain

Rationale: Intelligence analysis cohered source-evaluation matrices such as the Admiralty Code, separating source reliability from report credibility and updating both with performance.

Related originating lineages:

Review resolution: Security and intelligence is primary because source evaluation is a standing analytic tradecraft; statistics supplies calibration and evidence-weighting discipline.

Review outcome: Reconciled after independent review; high confidence.

Notes

[n1] The NATO "Admiralty Code" (or source-evaluation matrix) rates a source's reliability on one lettered scale (A–F) and each report's credibility on a separate numbered scale (1–6), deliberately keeping the two axes independent so a reliable source's dubious report is not automatically believed. It is the canonical model for a multi-axis, human-legible reliability scorecard. ↩a ↩b