Skip to content

Judgment Aggregation Dashboard

Metric / dashboard — instantiates Structured Expert Judgment Iteration

Displays distributions, movement between rounds, confidence, subgroup variation, and unresolved disagreements so iteration remains visible.

A Judgment Aggregation Dashboard is the display surface of the archetype: a live visual that renders the pooled expert judgments as spreads, round-over-round movement, subgroup fans, and still-open disagreements, so that iteration is something people can watch rather than infer. Its defining move is making the shape of the aggregate legible at a glance — not "the group says 40%" but "the group is bimodal, the epidemiologists cluster low, and the gap between clusters barely moved from round two to round three." Where a report is written once and a study runs the rounds, the dashboard is the always-current window onto them; it computes and shows, it does not elicit and does not decide. Its whole value is refusing to collapse a distribution into a headline mean.

Example

A regional infectious-disease team pools weekly forecasts from a dozen modeling groups for hospital admissions four weeks out, and the coordinators need to see the ensemble without flattening it — the way real forecast hubs present combined predictions.[1] A Judgment Aggregation Dashboard renders each group's predictive interval as a fan, overlays the pooled distribution, and — critically — shows how the fans shifted since last week's round, with a small panel breaking the pool out by model type (mechanistic vs. statistical).

The display earns its keep on a week when the naive average looks reassuringly stable. The dashboard shows the average is stable only because two clusters moved in opposite directions — the mechanistic models swung sharply upward on a new variant signal while the statistical models held flat — and the between-cluster gap widened. A single mean would have hidden a live disagreement about a possible surge; the dashboard puts it on screen, prompting the coordinators to treat the "stable" forecast as contested rather than settled.

How it works

  • Ingest and pool. Take the round's individual judgments and compute the aggregate as a distribution — quantiles, density, or a fan — never only a central tendency.
  • Show movement. Render this round against prior rounds so shifts, narrowing, and stubborn gaps are visible as change, not just as a new snapshot.
  • Break out subgroups. Split the pool by discipline, method, or region to reveal clusters a global aggregate hides.
  • Surface the unresolved. Highlight where dispersion stayed wide or bimodal, flagging live disagreement rather than letting it dissolve into the average.

All identities are stripped before display — the dashboard shows the shape of the group's views, never who holds them.

Tuning parameters

  • Aggregation rule — simple average, trimmed mean, or performance-weighted pool. Weighting can sharpen the signal but embeds a contestable judgment about whose view counts more.
  • Distribution rendering — interval, box, or full density. Fuller renderings preserve dispersion but can overwhelm a decision-maker who wanted a number.
  • Subgroup cuts — which partitions are shown. Revealing clusters aids insight but risks re-identifying small subgroups and eroding anonymity.
  • Movement window — how many prior rounds are overlaid. More history shows trajectory; too much clutters the current picture.

When it helps, and when it misleads

Its strength is keeping dispersion and disagreement visible through the whole process — it is the standing defense against overconfident aggregation, because a bimodal pool simply cannot be mistaken for a confident consensus once it is drawn on screen. It also makes iteration legible: you can see judgments converge, or stubbornly refuse to.

Its failure mode is the authority of a clean chart: a smooth aggregate curve can imply precision the underlying judgments do not have, and a well-designed dashboard can make a thin, poorly-sourced pool look authoritative. The classic misuse is letting the headline aggregate become the takeaway anyway — readers glance at the mean and ignore the fan around it. The guarding discipline is to render dispersion as prominently as the central estimate, keep subgroup cuts coarse enough to protect anonymity, and label the aggregate as a summary of judgment, not a measurement.

How it implements the components

  • uncertainty_distribution — it renders the pooled judgment as a spread or density, its foundational refusal to collapse to a mean.
  • convergence_disagreement_report — its movement and subgroup panels are a live, visual convergence/disagreement view, showing where the pool settled and where it split.
  • anonymized_feedback — it returns the aggregate to the panel between rounds with identities stripped, the reflected image experts revise against.

It displays; it does not elicit or terminate. It does not capture the raw judgments (independent_elicitation is Anonymous Survey Round's) and does not decide when to stop or record why views changed (stopping_rule is Delphi Study's; rationale_trace is Rationale Coding Matrix's).

Editorial Notes

Form Classification

Form family: Monitoring, Sensing & Alerting

Rationale: The dashboard repeatedly updates distributions, round-to-round movement, confidence, subgroup variation, and unresolved disagreement.

Nearest alternative: Interface, Display & Cue — The visible surface aids interpretation, but continuing observation of judgment state is primary.

Review outcome: Adjudicated after independent review; high confidence.

Origin Attribution

Primary origin: Operations Research

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Decision analysis developed structured expert-judgment aggregation and displays of dispersion, confidence, and disagreement.

Related originating lineages:

Review resolution: Both independent reviews place the primary lineage in operations_research. The queued differences (reported_ambiguity) concern secondary metadata rather than primary provenance. The final retains futurism_foresight, statistics_experimental_design only where a reviewer supplied a formative-lineage rationale; downstream application by itself is not treated as origin. origin_mode=cross_disciplinary_synthesis records the relationship among origin traditions, while domain_reach=multi_domain records application breadth separately. encyclopedia_synthesis=true reflects whether either reviewer identified a corpus-specific synthesis, and confidence=medium preserves the more cautious evidence assessment.

Attribution caveat: The dashboard packaging is modern, while its elicitation and aggregation components have several lineages.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Reconciled after independent review; medium confidence.

References

[1] Cramer, E. Y., et al. "Evaluation of individual and ensemble probabilistic forecasts of COVID-19 mortality in the United States". Proceedings of the National Academy of Sciences 119(15), e2113561119 (2022). Reports a real forecast hub that combined forecasts from participating models into an ensemble prediction. registry