Scoring Rule¶
Evaluate a probabilistic forecast after its outcome by mapping the report–outcome pair to a numeric loss or reward, with propriety governing whether truthful distributions are optimal in expectation.
Core Idea¶
A scoring rule is a numerical evaluation rule for probabilistic forecasts. It takes an issued probability distribution P and the outcome y that actually occurs, then returns a real-valued loss or reward S(P,y). The rule makes forecasts with different uncertainty shapes commensurable after the event while retaining the distributional claim that was made before it. A forecaster who said “rain probability 0.7” and one who said “0.5” are not judged only by whether rain occurred; their stated degrees of uncertainty determine different scores for the same realization.
The identity is not exhausted by a distance formula. A complete scoring-rule setup fixes an outcome space, an admissible class of forecast distributions, an orientation (lower loss is better or higher reward is better), integrability and boundary conventions, and a protocol for aggregating realized scores across cases. It also states whether the rule is proper. Under the loss orientation, define the expected score of report P when outcomes are actually generated from Q as S(P,Q)=E_{Y~Q}[S(P,Y)]. The rule is proper when S(Q,Q) <= S(P,Q) for every admissible P,Q: reporting the data-generating distribution is at least as good in expectation as any misreport. It is strictly proper when equality requires P=Q. Under the reward orientation, the inequalities reverse. Gneiting and Raftery give the general probability-space formulation and connect proper rules to entropy, divergence, estimation, and forecast comparison.[1]
Propriety is a design property of a subclass, not part of the bare definition of every scoring rule. An improper rule can still map forecasts and outcomes to numbers, but it creates an incentive to distort probabilistic reports or systematically rewards a feature other than the full predictive distribution. This distinction is the node's most important internal boundary. “Scoring rule” names the report–outcome evaluator; “proper scoring rule” adds an expected-score truthfulness condition; “strictly proper scoring rule” adds unique identification of the true distribution within the admitted class.
Two canonical families show the range. The quadratic or Brier loss for a binary event is S(p,y)=(p-y)^2, where p is the reported event probability and y is 0 or 1. Brier introduced the probability score for forecast verification in meteorology.[2] The logarithmic loss is S(P,y)=-log p(y), assigning a large penalty to an outcome given very small probability and infinite loss to an event assigned zero probability under the ideal mathematical form; Good linked this logarithmic score to rational probability assessment and information.[3] Both are strictly proper on appropriate domains, but they attend to different features and have different robustness and locality properties.
Structural Signature¶
Sig role-phrases:
- the outcome space — the mutually exclusive events, categories, values, vectors, paths, or other realizations that may occur
- the admissible forecast class — the probability distributions or densities that forecasters are permitted to report and relative to which validity is claimed
- the issued probabilistic report — a complete predictive distribution
P, fixed before the outcome is observed - the realized outcome — the observed
yagainst which the report is evaluated without retroactively changing the forecast - the report–outcome map — a rule
S(P,y)producing a numerical loss or reward from both inputs - the orientation and boundary conventions — lower-versus-higher-is-better, treatment of zero probabilities, infinite values, missing outcomes, and incomparable support
- the expected-score functional — expectation of the realized score under a candidate data-generating distribution
Q - the propriety status — proved proper, strictly proper, improper, or unestablished relative to the stated forecast class
- the aggregation and comparison protocol — the sampling unit, weights, averaging rule, uncertainty estimate, and permitted model or forecaster comparisons
The locked recognition test requires all core roles. A candidate is a scoring rule in this sense only if the reported object is a probability distribution, not merely a point; the outcome is realized after the report; a declared function uses both distribution and realization to produce a numeric loss or reward; and the domain, orientation, and aggregation semantics are recoverable. If truthfulness is claimed, the expected-score inequality must be proved or cited for the stated distribution class rather than inferred from a favorable example.
Diagnostics localize failure. If multiplying by -1 changes which forecast is called better, orientation has been mixed. If a forecaster improves expected score by sharpening or flattening a true distribution, the rule is improper on the tested class. If a log score explodes because rounded zeros were admitted, boundary handling is broken. If average scores reverse after changing case weights, the comparison is sampling-frame dependent. If forecasts with distinct calibration and sharpness profiles tie, the scalar score has compressed diagnostically relevant structure. If one dramatic tail miss dominates the result, determine whether this is the intended sensitivity of the rule or an unacknowledged robustness failure.
Interventions follow the diagnosis. State the orientation and admitted forecast class. Prove propriety or select a known proper rule appropriate to the outcome type. Clip or regularize probabilities only with the induced target and incentive distortion documented. Report score uncertainty and paired casewise differences rather than naked means. Add calibration and sharpness diagnostics when a scalar obscures failure structure. Use multiple prespecified proper rules when no single sensitivity profile matches every consequential aspect of the forecast.
What It Is Not¶
- Not a scoring function for a point forecast. A point-scoring function evaluates a reported mean, median, quantile, class, or other functional. A scoring rule evaluates the whole predictive distribution. Consistency for a functional is related to, but distinct from, propriety for distributions.
- Not a loss function in general. Optimization losses may score actions, parameters, residuals, or class margins. A scoring rule has the narrower forecast-distribution plus realized-outcome signature.
- Not forecast calibration. Calibration is a joint long-run agreement property of forecasts and outcomes. A proper score combines calibration and concentration into one performance measure; a single realized score cannot establish calibration.
- Not sharpness alone. A degenerate forecast can be maximally concentrated and catastrophically wrong. Sharpness is desirable only subject to calibration.[4]
- Not prediction error alone. A signed residual compares a point prediction with an observation. Distributional scores account for the entire uncertainty report; log loss, for example, reads the probability assigned to what occurred.
- Not expected utility. Expected utility ranks risky actions from probabilities and outcome utilities. Expected score evaluates a forecast report under a possible data-generating distribution and, for proper rules, tests whether truthful reporting is optimal.
- Not incentive compatibility automatically. Propriety gives a limited truth-telling incentive under an expected-score objective. An arbitrary or improper scoring rule lacks it, and real elicitation can introduce wealth, risk-attitude, strategic, or participation effects beyond the mathematical score.
- Not an accuracy rate. Thresholding probabilities into correct/incorrect labels discards uncertainty and usually creates an improper evaluation for full probabilistic reports.
- Not one named metric. Brier, logarithmic, spherical, continuous ranked probability, and energy scores are different scoring rules with different domains and sensitivities.
Scope of Application¶
Scoring rules apply wherever complete predictive distributions are issued and later confronted with realized outcomes.
- Binary event forecasting. Rain, recession, equipment failure, disease onset, and default probabilities can be evaluated with Brier or logarithmic scores while preserving confidence rather than reducing forecasts to yes/no calls.
- Multicategory prediction. A probability vector over mutually exclusive classes can be scored against the realized category. The rule evaluates allocation of probability across all categories, not only the winning label.
- Density forecasting. Continuous outcomes such as temperature, demand, wind speed, or asset returns require rules defined on densities or cumulative distributions. Matheson and Winkler developed proper-scoring families for continuous probability distributions.[5]
- Forecast comparison. Casewise scores from competing models can be averaged or otherwise aggregated on a common evaluation sample, with paired uncertainty analysis and predeclared weighting.
- Expert probability elicitation. Payments based on a proper rule can make truthful reporting optimal in expectation under the rule's assumptions. Savage analyzed proper scoring rules as devices for eliciting personal probabilities and expectations.[6]
- Probabilistic model estimation and training. Minimizing an empirical strictly proper score can fit predictive distributions; log loss is a familiar example. The statistical target depends on the rule and model class, so “training loss” and “forecast evaluation” should remain distinguished in protocol even when the formula is shared.
- Weather and climate verification. Probability forecasts are repeatedly issued and observed, enabling Brier-score comparison, reliability analysis, resolution analysis, and calibration/sharpness diagnostics.
- Risk and medical prediction. Scoring rules compare distributions or event probabilities, but clinical or financial usefulness can still depend on asymmetric downstream decisions not represented by the general score.
- Prediction markets and forecasting tournaments. Scores can rank participants or determine rewards, provided the elicitation mechanism, information timing, dependence among questions, and incentive assumptions are explicit.
The method's scope ends where the reported object is not probabilistic, the outcome is not well-defined, distributions place mass outside a common support, observations are selectively missing, or the evaluation sample is not representative of the claim. A valid formula cannot repair a broken target or observation protocol.
Clarity¶
The first clarity rule is to declare orientation. Literature uses both scores-as-rewards and scores-as-losses. In a reward convention, larger is better and propriety means truthful reporting maximizes expected score. In a loss convention, smaller is better and truth minimizes expected score. Multiplication by -1 converts conventions without changing rankings, but silently combining them reverses conclusions. This draft uses loss orientation unless explicitly noted.
The second rule is to separate realized score from expected score. A truthful 0.9 forecast for an event that genuinely occurs with probability 0.9 will sometimes receive a worse realized score than a timid 0.6 report when the rare non-event occurs. Propriety is not a promise that truth wins every realization; it is an expectation statement under repeated or hypothetical draws from the true distribution. Penalizing a forecaster for a single low-probability event as though it disproved the forecast confuses probability with certainty.
The third rule is to make the reference class explicit. Properness is relative to the admitted distributions and integrability conditions. A score can be strictly proper on one class and undefined or non-identifying on another. Log loss needs a density or mass at the realized outcome and creates boundary problems when reported probability is zero. Energy and kernel scores can identify distributions only under their own conditions. “Proper” without “for which class?” is incomplete.
The fourth rule is to distinguish distribution scoring from functional scoring. Squared error is a consistent scoring function for the conditional mean when the report is a point. The same quadratic form applied to a binary event probability is the Brier scoring rule for a Bernoulli distribution. Formula resemblance does not erase the different report objects or validity claims.
Finally, a scalar ranking does not diagnose why one forecaster wins. Proper scores reward calibration and sharpness jointly, but the same mean can arise from different error structures. Murphy's Brier-score partition separates uncertainty, reliability, and resolution components, making the source of performance visible.[7] Use the score for accountable comparison and a decomposition or graphical diagnostic for repair.
Manages Complexity¶
A predictive distribution is high-dimensional: it expresses location, spread, asymmetry, tails, multimodality, and dependence. A realized outcome supplies only one sample from that claim. A scoring rule compresses the report–realization relation into one number that can accumulate across forecast cases. Without such a common map, forecasters can selectively emphasize whichever aspect makes their forecast look favorable after the fact.
Properness stabilizes this compression. It rules out a class of score designs that would systematically invite hedging away from the forecaster's actual distributional assessment. Strict propriety adds identification: within the admitted class, no distinct report ties truth in expectation. This allows forecast generation, elicitation, model fitting, and evaluation to share a coherent target, though real-world incentives may require additional mechanism design.
Different rules compress differently. Log loss is local: it depends on the probability or density assigned to the outcome that occurred, and it strongly punishes small assigned probabilities. Quadratic/Brier loss responds smoothly to probability error for categorical outcomes. The continuous ranked probability score integrates squared differences between forecast and outcome cumulative distributions, producing a distance-sensitive assessment for ordered continuous outcomes. There is no context-free “best” sensitivity; selecting a rule is part of the evaluation design.
Aggregation turns casewise scores into a comparison, but introduces another layer of modeling. Equal weighting estimates performance over the empirical case mix; importance weights target another population; time averaging assumes a stable or explicitly handled dependence structure. Missing outcomes, clustered cases, and repeated forecasts from the same event alter uncertainty. A mean score without its estimand and uncertainty can be precise-looking noise.
The right management strategy is layered: preserve individual scores; aggregate under a declared sampling frame; quantify paired uncertainty; inspect calibration and sharpness; decompose where available; and keep the raw forecast–outcome records auditable. The scalar makes comparison possible, while the retained structure keeps the comparison interpretable.
Abstract Reasoning¶
Report–outcome separation. Freeze the report before observing the realization. Any score computed from a revised distribution answers a different question and permits hindsight leakage.
Expected-score audit. For candidate truth Q, compute E_Q S(P,Y) as a function of possible reports P. Check whether P=Q is a minimum under loss orientation and whether it is unique. This is the definitive propriety test, not the rule's name or intuitive appeal.
Divergence reading. For many strictly proper rules, excess expected loss S(P,Q)-S(Q,Q) is a nonnegative divergence that vanishes only at equality. This separates irreducible entropy or uncertainty under Q from regret caused by reporting P.
Orientation normalization. Before comparing formulas, convert all rules to a common higher- or lower-is-better convention. Preserve affine transformations only when they do not vary across forecasts or outcomes in a way that changes the target.
Boundary stress test. Evaluate zero and near-zero assigned probabilities, heavy tails, support mismatch, and missing observations. A mathematically proper rule can be operationally unusable if recording precision creates infinite or dominant penalties.
Sensitivity profile. Ask which changes in a forecast the score notices most. Log loss focuses on assigned density at the realization; quadratic rules spread sensitivity across categories; distance-sensitive rules use outcome geometry. Choose deliberately rather than treating all proper rules as interchangeable.
Sampling-frame audit. Define the population of forecast cases and outcome-verification process. Use paired differences when competing forecasts cover the same cases, and avoid comparing means from different case mixes without adjustment.
Diagnostic decomposition. After ranking, inspect reliability/calibration, resolution/sharpness, subgroup performance, and tail contributions. A ranking answers “which scored better under this protocol”; decomposition answers “why and where.”
Knowledge Transfer¶
Within forecast evaluation, the same package transfers from rain probabilities to disease risks, credit defaults, demand densities, categorical classifiers, and multivariate trajectories. The outcome space and scoring family change, but the audit questions persist: What distribution was reported? When was it frozen? Which outcome occurred? What rule and orientation were used? Is propriety established for the class? How were cases weighted? What uncertainty surrounds the aggregate?
The package also transfers from evaluation to elicitation and training with controlled changes in interpretation. As an evaluation tool, a score compares probabilistic forecasts after outcomes. As an elicitation payment, expected score influences what a forecaster reports. As a training objective, empirical score guides parameter fitting. The same formula can occupy all three roles, but dependencies, strategic behavior, model misspecification, and overfitting differ. Labeling the role prevents evidence from one use being imported uncritically into another.
The distinction between rule and propriety supports cumulative design. New scores can be classified by report domain, locality, robustness, sensitivity, and truthfulness. A method developer need not reinvent forecast-evaluation logic; they must demonstrate where the new rule sits in this matrix and prove or delimit its target.
Outside probabilistic forecasting, “score what was claimed after observing what happened” resembles performance evaluation broadly. The literal scoring-rule identity does not travel unless the report is a probability distribution and propriety is defined through expectation under a distribution. Generic transfer is already owned by evaluation, comparison, expected_utility, and incentive_compatibility.
Examples¶
Canonical: Why the binary Brier loss rewards the true probability¶
Let a binary outcome Y equal 1 with true probability q=0.7. A forecaster reports p, and the loss is S(p,Y)=(p-Y)^2. The expected loss is
E[S(p,Y)] = q(1-p)^2 + (1-q)p^2 = (p-q)^2 + q(1-q).
The second term, q(1-q)=0.21, does not depend on the report. The first term is nonnegative and vanishes only when p=q. Reporting p=0.7 therefore yields expected loss 0.21; reporting p=0.5 yields 0.25; reporting p=0.9 also yields 0.25. The true probability is the unique optimum, so binary Brier loss is strictly proper on this report space. One realized non-event would give the truthful report loss 0.49 and the timid report loss 0.25, illustrating why realized and expected comparisons must not be confused.[2][1]
Mapped back: The outcome space is {0,1}; the admissible reports are Bernoulli distributions indexed by p; the issued report is frozen before Y; the report–outcome map is squared loss; lower is better; expectation is taken under Bernoulli q; the unique minimum proves strict propriety; and repeated cases provide the aggregation protocol.
Applied / In Practice: Compare and recalibrate probabilistic rain forecasts¶
Consider 100 daily forecasts. On 50 days a model reports rain probability 0.8, and rain occurs on 35; on 50 days it reports 0.2, and rain occurs on 15. Its total binary Brier loss is 35(0.8-1)^2 + 15(0.8-0)^2 + 15(0.2-1)^2 + 35(0.2-0)^2 = 22, for a mean of 0.22. A constant 0.5 forecast has mean loss 0.25, so the model's separation of higher- and lower-risk days has value.
The grouped outcomes also reveal miscalibration: events forecast at 0.8 occur at frequency 0.7, and events forecast at 0.2 occur at frequency 0.3. Replacing group reports with 0.7 and 0.3, if justified on held-out data, reduces the sample mean Brier loss to 0.21. The score ranks the protocols; reliability/resolution decomposition and calibration plots diagnose the overconfidence and show where recalibration acts.[7][4]
Mapped back: The daily rain event supplies the outcome space and realization; each probability is an issued Bernoulli distribution; Brier loss maps every report–outcome pair to a lower-is-better number; equal weighting defines the sample mean; the baseline comparison reads forecast value; and the reliability diagnostic supplies a targeted intervention without changing the rule's identity.
Structural Tensions¶
Truthfulness in expectation ↔ realized-outcome volatility. A proper rule rewards truth over repeated or hypothetical draws, yet a rare event can make an honest forecast look bad on one case. Diagnostic: compare expected-score logic and confidence intervals, not just the most salient realization.
Calibration ↔ sharpness. Broad distributions can be calibrated but uninformative; concentrated distributions can be useful but overconfident. Diagnostic: inspect calibration and concentration separately alongside the proper score.
Tail sensitivity ↔ robustness. Log loss appropriately punishes assigning tiny probability to what occurs, but one rounded zero or contaminated observation can dominate an aggregate. Diagnostic: plot casewise contributions, audit support and precision, and prespecify any clipping.
Locality ↔ distributional sensitivity. A local score reads density at the realized outcome; a nonlocal score can respond to the rest of the predictive distribution. Diagnostic: perturb unobserved regions of P and record whether and why the score should change.
Universal comparison ↔ decision relevance. A general proper rule gives an application-neutral ranking, while a downstream decision may care asymmetrically about particular thresholds or tails. Diagnostic: separate distributional accuracy from a declared decision-specific utility analysis.
Scalar compression ↔ diagnostic resolution. One mean score supports ranking but can hide calibration, resolution, subgroup, time, and tail failures. Diagnostic: retain decompositions, subgroup profiles, and paired casewise differences.
Stable protocol ↔ changing case mix. Fixed weights permit comparability, yet shifts in prevalence, horizon, or verification can reverse performance. Diagnostic: declare the target case distribution and rerun standardized or stratified comparisons under relevant shifts.
Method autonomy ↔ reduction to generic evaluation. A scoring rule is a formal evaluation, but reducing it to “apply a criterion and output a number” erases the probabilistic report, realized random outcome, expected-score functional, and propriety condition. Diagnostic: replace distributions with arbitrary objects; if expected truth-telling remains definable without reconstructing probability semantics, the account is too generic.
Structural–Framed Character¶
Scoring Rule is structural-leaning with an aggregate framedness score of 0.2. It is a formal, mechanizable evaluation instrument whose identity nevertheless remains native to probabilistic forecasting and statistical decision theory.
- Vocabulary travels (0.5). “Rule,” “score,” “report,” and “outcome” travel widely. “Predictive distribution,” “proper,” and “strictly proper” carry specialized probabilistic meanings.
- Evaluative weight (0.0). The score orders forecast performance by a declared mathematical criterion; it does not import a social or moral judgment about the forecaster.
- Institutional origin (0.0). Tournaments and agencies use scores, but no institution constitutes the mathematical report–outcome mapping or propriety proof.
- Human-practice bound (0.0). Automated forecasts, outcomes, and scoring pipelines instantiate the rule without casewise human interpretation.
- Import versus recognize (0.5). The outcome is observed, but the scoring rule imposes a chosen sensitivity and orientation on how the forecast is evaluated.
Its character: a formal probabilistic evaluator with a constructed criterion and an objective, distribution-relative truthfulness test.
Structural Core vs. Domain Accent¶
Structural core. A claim is frozen before feedback; a realized state arrives; a declared function maps claim and state to a common scalar; expected performance evaluates systematic reporting strategies; and aggregated results support comparison and repair.
Domain accent. The claim is a probability distribution; the state is a draw from an outcome space; expectation is taken under a candidate true distribution; propriety means truthful probabilistic reporting is optimal; and rule validity depends on forecast class, support, integrability, and sampling protocol.
Why this does not clear the prime bar. Generic post-outcome evaluation appears across education, games, engineering, and governance, but those activities do not literally instantiate propriety over probability distributions. Removing the domain accent yields the live evaluation prime plus related expected-value and incentive structures. Retaining it yields a coherent method family used across forecast domains, not a substrate-independent prime.
Instantiates / Related Primes¶
evaluation. A scoring rule applies a criterion-bearing map to a bounded forecast and realized evidence, producing an interpretable score. This is the proposed strict subsumption parent.expected_utility. Both take expectations of outcome-dependent values. Expected utility ranks risky actions by outcome utility; expected score tests or ranks forecast reports by a score function.incentive_compatibility. Properness creates a limited truth-telling alignment when the forecaster optimizes expected score. The relationship applies to proper rules under behavioral assumptions, not to the whole scoring-rule genus.calibration. Calibration is a joint forecast–outcome property that proper scores help reward and diagnose. It is neither a synonym nor sufficient for sharp probabilistic forecasts.comparison. Aggregated scores place forecasters under a shared performance frame. A scoring rule can evaluate one forecast, so comparison is a common use rather than a constitutive requirement.measurement. The numeric output behaves like a standardized reading, but it is a criterion-dependent evaluation, not direct measurement of a target attribute with unit and calibration chain.prediction_error. Some rules transform a forecast–outcome discrepancy, but distributional scores need not be signed residuals and preserve more than a point prediction.
Relationships to Other Abstractions¶
Current abstraction Scoring Rule Domain-specific
Parents (1) — more general patterns this builds on
-
Scoring Rule is a kind of Evaluation Prime
evaluation. A scoring rule applies a criterion-bearing map to a bounded forecast and realized evidence, producing an interpretable score.This is the proposed strict subsumption parent.
Hierarchy path (1) — routes to 1 parentless root
- Scoring Rule → Evaluation → Comparison → Self Checking
Neighborhood in Abstraction Space¶
Scoring Rule sits in a sparse region of the domain-specific corpus (79th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Expectation, Retrospection & Evaluation Bias (12 abstractions)
Nearest neighbors
- Probability Distribution — 0.84
- Credal Set — 0.83
- Learnable Function Class — 0.82
- Random Variable — 0.82
- Kaplan–Meier estimator — 0.82
Computed from structural-signature embeddings · 2026-09-08
Not to Be Confused With¶
- Proper scoring rule. Tell: a scoring rule becomes proper only after an expected-score truthfulness inequality is established for a stated forecast class; propriety is a subclass property, not an alias for the genus.
- Scoring function for a point forecast. Tell: if the report is a mean, median, quantile, or class rather than a full distribution, the relevant property is consistency for that functional.
- Generic loss function. Tell: ask whether the first argument is a probabilistic forecast and the second a subsequently realized outcome. If not, the broader optimization-loss identity is cleaner.
- Calibration metric. Tell: calibration requires a collection of forecast–outcome pairs and tests statistical agreement; one scoring-rule value is defined for one pair.
- Accuracy or zero–one loss. Tell: if probabilities are thresholded before evaluation, confidence information has been discarded and full-distribution propriety is generally lost.
- Expected utility. Tell: expected utility chooses among actions using utilities of consequences; scoring rules evaluate distributions that were reported.
- Probability weighting function. Tell: probability weighting transforms probabilities to model decision behavior; a scoring rule consumes the reported distribution and realization to evaluate forecasting performance.
- Brier score. Tell: Brier is one quadratic member of the scoring-rule family, not the family itself.
- Log loss / cross-entropy. Tell: logarithmic loss is one strictly proper rule under suitable conditions; its locality and boundary behavior do not define every scoring rule.
- Game or exam scoring rule. Tell: if the score depends on points, votes, or answers without a complete probabilistic report and an expected-score propriety question, it is a different sense of “scoring rule.”
References¶
[1] Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378. https://doi.org/10.1198/016214506000001437 registry ↩a ↩b
[2] Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3. https://doi.org/10.1175/1520-0493(1950)078%3C0001:VOFEIT%3E2.0.CO;2 registry ↩a ↩b
[3] Good, I. J. (1952). Rational decisions. Journal of the Royal Statistical Society: Series B, 14(1), 107–114. https://doi.org/10.1111/j.2517-6161.1952.tb00104.x registry ↩
[4] Gneiting, T., Balabdaoui, F., & Raftery, A. E. (2007). Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society: Series B, 69(2), 243–268. https://doi.org/10.1111/j.1467-9868.2007.00587.x registry ↩a ↩b
[5] Matheson, J. E., & Winkler, R. L. (1976). Scoring rules for continuous probability distributions. Management Science, 22(10), 1087–1096. https://doi.org/10.1287/mnsc.22.10.1087 registry ↩
[6] Savage, L. J. (1971). Elicitation of personal probabilities and expectations. Journal of the American Statistical Association, 66(336), 783–801. https://doi.org/10.1080/01621459.1971.10482346 registry ↩
[7] Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4), 595–600. https://doi.org/10.1175/1520-0450(1973)012%3C0595:ANVPOT%3E2.0.CO;2 registry ↩a ↩b