Skip to content

Consequence Design Review

Protocol — instantiates Reinforcement Loop Design

Reviews proposed rewards, penalties, feedback, recognition, and natural consequences for alignment, proportionality, timing, fairness, and side effects.

Version
v1 · 2026-08-24 · History
Mechanism #
1798
Type
Protocol
Form family
Assessment, Review & Assurance
Solution family
Learning & Scaffolding
Problem family
Learning, Knowledge & Capability Gaps
Problem subfamily
Adaptive Feedback, Reinforcement & Calibration
Origin domain
Psychology
Also from
Law & Governance, Ethics of Technology & AI Governance
Instantiates
Reinforcement Loop Design

Before a consequence is switched on for real people, someone should ask whether it is aimed at the right thing, sized right, and allowed to touch the people it touches. Consequence Design Review is that deliberative gate: a structured review that takes a proposed reward or penalty and checks it against the behavior it is meant to serve, the proportionality of its magnitude, and the ethical limits on how it may cue, observe, or coerce. Its defining move is that it is evaluative and normative, not generative — it does not invent exploits and it does not build the reward; it judges a design against a checklist of alignment, proportionality, and consent, and sends it back to be fixed or approves it to ship. It is the loop's ethics-and-alignment review board.

Example

A food-delivery platform proposes a new consequence structure for couriers: a star rating from customers, an automatic penalty (reduced order priority) below 4.6 stars, and deactivation below 4.2. Before rollout it goes to a Consequence Design Review. The review works down its checklist. Alignment: the stated goal is reliable, safe delivery — but a rating driven largely by factors outside the courier's control (restaurant delays, weather) is only loosely tied to that goal, so the consequence is aimed at a noisy proxy. Proportionality: deactivation — someone's livelihood — for a 0.4-star swing that a handful of unfair reviews can cause is grossly out of scale with the behavior it punishes. Consent and dignity: the rating is opaque; couriers can't see or contest the reviews that end their income.

The review does not redesign the whole scheme, but it returns three binding conditions: re-anchor the rating on controllable behaviors, add a proportionate graduated response with a human appeal before deactivation, and make the rating transparent and contestable. Only then does the consequence clear the gate. The review built nothing and blocked nothing outright; it held the proposed consequence to alignment, proportionality, and consent — and made those the price of shipping.

How it works

  • Check alignment against the real goal. Confirm the consequence reinforces the actual behavior of interest, not a convenient proxy that can improve while the goal decays.
  • Test proportionality. Weigh the magnitude of the reward or penalty against the behavior — over-strong consequences distort judgment and invite gaming; life-altering penalties demand especially high scrutiny and due process.
  • Set the consent and dignity boundary. Decide what forms of cueing, observation, personalization, and penalty are acceptable for these people in this setting, and require transparency and contestability where stakes are high.
  • Weigh side effects and consume the red team's findings. Fold in the exploit list from adversarial testing and judge whether the design's residual risks are tolerable, then approve, condition, or reject.

Tuning parameters

  • Review rigor — a lightweight checklist versus a full panel with appeal design. Heavier review catches more harm but slows deployment; match it to the stakes of the consequence.
  • Proportionality bar — how tightly penalty magnitude must track behavior severity. A strict bar protects people but blunts deterrence; a loose one is punchy and dangerous.
  • Consent threshold — how much transparency and opt-out is required before a consequence may touch someone. Higher thresholds protect autonomy but constrain design; lower ones enable manipulation.
  • Due-process depth — whether high-stakes penalties require human appeal. Appeals protect against unfair automated harm at a cost in speed and overhead.
  • Scope of review — the consequence in isolation versus its interaction with existing incentives. Wider scope catches compound harms but lengthens the review.

When it helps, and when it misleads

Its strength is that it is the loop's conscience before launch: it catches the disproportionate penalty, the misaligned reward, and the non-consensual nudge while they are still on paper, and it forces due process onto consequences that can seriously harm people. Its clearest job is keeping a loop on the right side of the line between a legitimate incentive and a dark pattern — a design that steers people toward outcomes serving the designer rather than themselves.[n1]

Its failure mode is that a review is only as honest as its checklist and its independence: run by the same people who built the loop, it becomes a rubber stamp that launders a bad design with a compliance signature. It can also over-correct into a bottleneck that reviews trivial nudges as heavily as livelihood-altering penalties, teaching teams to route around it. And because it is deliberative, it can miss the concrete exploit an adversarial test would have surfaced — which is why it should consume red-team findings rather than substitute for them. The guarding discipline is to keep the reviewer independent of the designer, to scale scrutiny to stakes, and to treat the red team's exploit list as a required input rather than trusting a clean-looking design.

How it implements the components

Consequence Design Review realizes the alignment-and-limits side of the loop — the components that judge a proposed consequence, none that build or schedule it:

  • behavior_goal — it holds the true target the consequence must serve, and rejects designs aimed at a convenient proxy instead.
  • autonomy_and_consent_boundary — it sets the ethical limits on cueing, observation, penalty, and personalization, and requires transparency and due process at high stakes.
  • reward_calibration — it judges the proportionality of the proposed reward or penalty magnitude (calibration the Reward or Recognition System proposes and this review vets).

It builds no consequence and runs no exploit hunt. The actual reward or penalty (consequence) is supplied by Reward or Recognition System, Immediate Feedback Interface, or Safety Reinforcement Protocol; the adversarial perverse_incentive_check is generated by Perverse Incentive Red Team. Among its protocol twins, it sets no reinforcement_schedule or fade_or_transfer_plan (Reinforcement Schedule Design) and installs no safe replacement_response (Safety Reinforcement Protocol).

Editorial Notes

Form Classification

Form family: Assessment, Review & Assurance

Rationale: Reviews proposed rewards, penalties, feedback, recognition, and natural consequences for alignment, proportionality, timing, fairness, and side effects, making its operative form a bounded evaluation of existing evidence or work that produces a finding or disposition.

Independent corroboration: The frozen evidence defines Consequence Design Review as 'Reviews proposed rewards, penalties, feedback, recognition, and natural consequences for alignment, proportionality, timing, fairness, and side effects', so its operative form is Assessment, Review & Assurance.

Review outcome: Independent reviewer agreement; high confidence.

Origin Attribution

Primary origin: Psychology

Origin pattern: Cross-disciplinary synthesis

Present-day reach: Multi-domain

Rationale: Behavior analysis established the design and ethical review of reinforcement and punishment consequences; technology governance and law supplied predeployment impact review, rights, proportionality, transparency, and appeal constraints.

Related originating lineages:

  • Law & Governance — Due-process and proportionality traditions contribute fairness, contestability, and dignity constraints on penalties.
  • Ethics of Technology & AI Governance — Technology governance supplied structured impact assessment, affected-community review, documentation, and predeployment go/no-go decisions.

Review resolution: The BACB ethics code requires behavior-change interventions, including consequences, to prioritize reinforcement, consider risks, benefits and side effects, minimize harm, and undergo review before implementation. NIST's AI RMF independently formalizes impact assessment, affected-community input, and predeployment go/no-go decisions. Psychology therefore supplies the consequence-design core, with technology governance and law retained for impact review, proportionality, consent, and due-process constraints.

Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.

Review outcome: Researched adjudication after independent review; high confidence.

Sources consulted:

Notes

[n1] Dark pattern — a term coined by Harry Brignull for interface and incentive designs that steer people into choices that serve the designer at the user's expense. Keeping a reinforcement loop on the legitimate side of that line is the sharpest test this review exists to apply.