Success Metric Reweighting¶
Scorecard policy — instantiates Bycatch-Aware Selective Intervention Design
Rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
Success Metric Reweighting changes the objective the whole selective process is optimized against, so that success stops meaning "target captured" and starts meaning "target captured net of bycatch." Its defining move is structural: it relocates off-target harm from outside the metric — an externality the system is free to ignore — to inside it, as a term that subtracts from the score, and it attaches that term to a named owner accountable for it. Where the other mechanisms detect, gate, or repair bycatch, this one alters the incentive that made bycatch invisible in the first place. It is the archetype's structural cause addressed head-on.
Example¶
A regional fishery has for years scored vessels and its whole management plan on landed target catch — tonnes of the species they mean to sell. Discarded bycatch (undersized fish, non-target species, the occasional protected animal) is thrown back and never enters the scorecard, so a "record season" can coincide with heavy collateral mortality. Success Metric Reweighting rewrites the objective: the score now credits landed target catch and debits a weighted bycatch term, so a vessel that lands 5% more target but triples its discards ends up scoring worse, not better.
Setting the weight is the crux — how many tonnes of target catch one protected-species interaction is allowed to "cost" — and the reweighting names an accountable owner (a fisheries manager, not a diffuse "the fleet") answerable for that debit. Almost immediately the fleet's behavior shifts toward the gear and grounds that keep the debit low, because for the first time the number they are judged on moves with the harm they cause.
How it works¶
- Add an explicit bycatch term to the objective, sourced from the harm ledger, so the score reads target and non-target together instead of target alone.
- Set the exchange rate — the weight — that says how much target yield a unit of bycatch is allowed to offset. This weight is the value judgment, now made explicit instead of implied by the metric's silence.
- Attach the term to a named owner accountable for it, so the number has a person behind it rather than dissolving across everyone and therefore no one.
- Wire the reweighted metric into the scorecards, incentives, and reviews that actually drive behavior — an expanded metric nobody is measured on changes nothing.
Tuning parameters¶
- Bycatch weight — how heavily off-target harm is penalized against target yield. Too light and the term is cosmetic; too heavy and target capture is strangled. This dial is the policy.
- Term structure — a linear penalty, a steep convex penalty, or a hard cap folded into the score. Convex penalties bite hardest on the worst offenders.
- Harm resolution — one lumped bycatch number versus separately weighted classes, so a protected species can be scored far above a common one.
- Ownership scope — how narrowly the accountable owner is drawn. A specific owner concentrates responsibility but invites scapegoating; a diffuse one dilutes it into nothing.
- Revision cadence — how often the weights are revisited as values, evidence, and gaming tactics change.
When it helps, and when it misleads¶
Its strength is that it attacks the archetype's structural cause directly — a selector over-catches largely because the scorecard rewarded only target yield — so a single reweighting can shift behavior across the whole system without touching the selector's mechanics. Making the exchange rate explicit also drags a hidden value judgment (how much collateral is acceptable) into the open, where it can be argued and owned rather than smuggled in by omission.
Its failure modes are those of any metric that becomes a target. Once bycatch is scored, effort flows to the measured bycatch, and unmeasured or reclassified harm can grow in its shadow — the reweighted number improves while the world does not, the textbook expression of Goodhart's law.[n1] A too-light weight launders harm as "accounted for" while changing nothing; a mis-set exchange rate can trade away a rare irreversible harm for cheap target gains. And a metric with no accountable owner is inert. The classic misuse is reweighting after the fact, to make an existing practice score well rather than to change it. The discipline that guards against this is to weight from the harm's real severity rather than its convenience, keep the classes disaggregated so severe rare harms can't be averaged away, and pair the metric with an independent audit that checks whether measured bycatch still tracks real bycatch.
How it implements the components¶
expanded_success_metric— the mechanism is this metric: the objective rewritten to net a weighted bycatch term against target yield, so success and collateral harm are scored on one scale.accountable_harm_owner— it names the person answerable for the bycatch term, without whom the reweighted number is one nobody owns and therefore nobody defends.
It sets how bycatch is scored but not the hard line at which operations must stop — that threshold-as-halt is Bycatch Tolerance Stop Rule — nor does it verify the ledger it reads: that check is False-Capture Audit. It repairs nothing; turning a harm reading into remediation is Compensation and Restoration Trigger.
Related¶
- Instantiates: Bycatch-Aware Selective Intervention Design — reweighting supplies the expanded objective the whole appraisal optimizes against.
- Consumes: Non-Target Sentinel Sampling and the Bycatch Rate Dashboard supply the harm ledger the bycatch term is computed from.
- Sibling mechanisms: Non-Target Sentinel Sampling · Non-Target Impact Pre-Mortem · Selectivity Window Test · Selector Retuning Cycle · Bycatch Rate Dashboard · Bycatch Tolerance Stop Rule · Compensation and Restoration Trigger · Escape Hatch or Release Protocol · False-Capture Audit · Negative Filter or Exclusion Device
Editorial Notes¶
Form Classification¶
Form family: Rule, Policy & Commitment
Rationale: Success Metric Reweighting operates as a standing rule, threshold, contractual commitment, or policy constraint governing future conduct because it rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
Independent corroboration: The frozen evidence defines Success Metric Reweighting as 'Rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term', so its operative form is Rule, Policy & Commitment.
Nearest alternative: Decision, Gate & Allocation — Success Metric Reweighting includes features of a case-specific gate, selection, routing, prioritization, or resource disposition, but its defining operation is a standing rule, threshold, contractual commitment, or policy constraint governing future conduct.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Operations Research
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Universal
Rationale: Revising weights on multiple success criteria after evidence shows distortion is multi-criteria decision analysis under governance. DOE MCDA guidance formalizes criteria, weights, sensitivity, and tradeoffs; management determines legitimacy and statistics checks measurement behavior.
Related originating lineages:
- Economics & Finance — Externality pricing makes off-target costs subtract from benefit.
- Environmental Science & Climate Studies — Bycatch and impact accounting inspired explicit harm terms.
- Organizational & Management Science — organizational_management contributes organizational design, management, and operational governance to this mechanism's defining operation—Rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term—without displacing the selected primary historical lineage.
- Statistics & Experimental Design — Statistics, experimental design, and measurement theory supplies a parallel or contributing lineage for the mechanism's defining operation: rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
- Systems Thinking & Cybernetics — Systems thinking, feedback control, and cybernetics supplies a parallel or contributing lineage for the mechanism's defining operation: rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
- Ethics of Technology & AI Governance — Technology ethics and ai governance supplies a parallel or contributing lineage for the mechanism's defining operation: rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
Review resolution: The blind reviewers disagree on primary lineage (operations_research versus organizational_management). Authoritative or primary research supports operations_research as the best historical origin: Revising weights on multiple success criteria after evidence shows distortion is multi-criteria decision analysis under governance. DOE MCDA guidance formalizes criteria, weights, sensitivity, and tradeoffs; management determines legitimacy and statistics checks measurement behavior. The cited U.S. Department of Energy, Multi-Criteria Decision Analysis directly supports the mechanism's defining operation. All independently supported contributing domains are retained without an arbitrary cap. origin_mode=cross_disciplinary_synthesis records lineage, while domain_reach=universal records later applicability separately from provenance.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Researched adjudication after independent review; high confidence.
Sources consulted:
Notes¶
The bycatch weight is a values choice wearing the costume of a technical parameter. Handed to analysts to "calibrate," it quietly becomes a policy decision made by whoever tunes the number — how much collateral harm is acceptable — without anyone having decided it as policy. Keeping the weight visible and owned at the level that should own the trade-off is what stops the scorecard from legislating ethics by default.
[n1] Goodhart's law — "when a measure becomes a target, it ceases to be a good measure." Once a bycatch term is scored, effort can flow to improving the number itself (reclassifying, shifting harm to unmeasured classes) rather than reducing real harm, which is why a reweighted metric needs an independent check that the measure still tracks the thing it stands for. ↩