Pre-Registered Benchmark Policy¶
Protocol — instantiates Risk-Adjustment and Benchmark Selection
A protocol that fixes benchmark-selection rules before outcomes are evaluated.
The most corrosive way to cheat a performance evaluation is to choose the benchmark after seeing the result, then present the flattering comparison as if it had been the obvious one all along. Pre-Registered Benchmark Policy removes that freedom by ordering time. It is a protocol that writes down — before any outcome is known — exactly what claim is being tested, over what horizon, and which benchmark, factor set, and reference universe will be used to judge it, along with rules for what would count as success or abnormal performance. Once registered, the comparator is locked; any later deviation must be declared, justified, and marked as post-hoc. Its defining property is temporal precommitment: it does not build the benchmark or estimate the residual, it binds the choice of them to a moment before the data can whisper which choice would look best. It is the honesty scaffold the rest of the machinery hangs on.
Example¶
A philanthropic foundation is about to fund a two-year pilot of a job-training program and wants a credible read on whether it works. Everyone involved knows the temptation: if the raw numbers come in soft, someone will propose comparing against an easier baseline; if they come in strong, someone will pick a comparison that makes them look heroic. So before a single participant enrolls, the foundation registers a Benchmark Policy. It fixes the claim precisely — "does completion-to-employment at twelve months exceed the matched-comparison rate?" — names the evaluation horizon, specifies the comparison group construction and the risk factors it will adjust for, and states the threshold that would count as a real effect.
The document is time-stamped and shared with the external evaluator. Two years later the results arrive mixed. A program officer floats switching to a rosier regional-average benchmark, but the policy is on record: that switch is now a visible, justified deviation, flagged as post-hoc, not a silent swap. The pre-registered comparison stands as the primary claim. The policy did not make the program better or worse — it made the eventual judgment believable, because the yardstick was chosen when no one yet knew which yardstick would flatter.
How it works¶
- Fix the claim and horizon first. State precisely what is being evaluated, over what period, with what outcome measure — before any outcome is observed.
- Lock the comparator rules. Register the benchmark, factor set, reference universe, and adjustment logic that will be used, plus the success or abnormal-performance threshold.
- Time-stamp and escrow. Record the policy with a date and, ideally, an independent party, so its precedence over the results is verifiable.
- Mark every later change. Permit deviations, but require each to be declared, justified, and labeled post-hoc, keeping the pre-committed comparison as the primary claim.
Tuning parameters¶
- Specification completeness — how fully the benchmark is pinned down in advance. Fully specifying removes discretion but can lock in a choice that later proves ill-suited; leaving decision rules (not values) allows principled adaptation without reopening the shopping.
- Amendment protocol — how deviations are permitted and recorded. Strict protocols maximize credibility but can force use of a stale benchmark after a regime shift.
- Escrow independence — self-registration versus an external registry or auditor; more independence is more credible but heavier.
- Scope of precommitment — whether only the primary benchmark is fixed or the full grid of secondary analyses too.
- Horizon rigidity — how firmly the evaluation window is fixed against the temptation to stop early on a favorable reading.
When it helps, and when it misleads¶
Its strength is that it structurally defeats post-hoc benchmark shopping and the garden of forking paths — the quiet manipulation where a comparator chosen after the fact is passed off as principled — by making the timing of the choice auditable.[n1] It is what lets a later residual claim be read as a test rather than a rationalization.
Its failure mode is rigidity: a benchmark that was reasonable at registration can be invalidated by a regime shift, and a policy with no principled amendment path forces a stale or wrong comparison to be honored. Precommitment can also be theater — a vaguely worded policy leaves so much discretion that nothing is really bound, or amendments are so freely granted that the lock is nominal. The guarding discipline is to specify decision rules precisely enough to bind while allowing declared, justified adaptation, and to keep the amendment log as visible as the original — so precommitment constrains manipulation without ossifying into a comparison everyone knows is wrong.
How it implements the components¶
Pre-Registered Benchmark Policy fills the precommitment side of the archetype — the machinery that fixes choices in time before outcomes can bias them:
pre_analysis_benchmark_precommitment— it is the registered, time-stamped lock on the benchmark, factor set, and reference universe, with an amendment log for any deviation.claim_scope_and_evaluation_horizon— it pins the precise claim, outcome measure, and evaluation window up front, so the comparator is chosen to fit a fixed question rather than a known answer.
It does not construct the comparator (benchmark_construction_rule) — that is Style-, Sector-, or Case-Matched Benchmark — nor estimate the residual it later governs (risk_adjustment_mapping), which is Multi-Factor Performance Model.
Related¶
- Instantiates: Risk-Adjustment and Benchmark Selection — it supplies the precommitment discipline that keeps the whole benchmark choice honest.
- Sibling mechanisms: Multi-Factor Performance Model · Style-, Sector-, or Case-Matched Benchmark · Benchmark Attribution Report · Alternative-Benchmark Sensitivity Grid · Out-of-Sample Benchmark Validation · Case-Mix Risk Stratification Table
Editorial Notes¶
Form Classification¶
Form family: Rule, Policy & Commitment
Rationale: The mechanism fixes benchmark, factor, reference-universe, adjustment, horizon, and success rules before outcomes are observed.
Nearest alternative: Protocol, Workflow & Routine — Registration follows steps, but the deployed result is the persistent precommitted comparator policy.
Review outcome: Adjudicated after independent review; high confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Multi-domain
Rationale: Fixing benchmark selection before outcomes are known is a preregistration control from experimental design.
Related originating lineages:
- Economics & Finance — Finance contributes risk-adjusted benchmarking and the incentive to choose flattering comparators.
Review resolution: Light authoritative-source research resolves the primary-origin disagreement in favor of statistics experimental design. National Academies/NCBI: Toolkit for Fostering Open Science Practices directly documents the defining practice or theory described in the selected origin rationale. Other domains are retained only where the blind reviews identify material co-development or translation; broad application is recorded separately as domain_reach=multi_domain, while origin_mode=cross_disciplinary_synthesis describes the relationship among origin lineages.
Attribution caveat: The boundary with economics finance is substantive because that tradition materially developed or translated part of the mechanism; the cited provenance places the defining form in statistics experimental design.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
Notes¶
[n1] Pre-registration — filing a study's hypotheses and analysis plan before collecting or seeing outcomes — is the standard scientific guard against outcome-contingent analytic choices; a pre-registered benchmark policy applies the same logic to comparator selection. ↩