Salience Normalization Test¶
Test or assessment — instantiates Supernormal Cue Guardrail Design
Compares decisions made under amplified versus normalized cue presentation to measure whether a choice depends on exaggerated salience rather than substantive value.
How do you prove a cue is doing the work instead of the underlying value? Salience Normalization Test answers empirically: present the same substantive choice twice — once with the cue amplified as designed, once with it normalized to a calibrated baseline — and measure whether the decision changes. Its defining idea is the controlled amplified-versus-normalized comparison of behavior: it does not inspect the design, it runs an experiment on responses and reads off cue-dependence from the difference. If choices are stable across presentations, the cue is riding on real value and is legitimate. If choices swing when only the salience is dialed down — content held constant — the response depends on the exaggeration, and the cue is a hijack candidate. It converts a suspicion ("this feels manipulative") into a measured effect size.
Example¶
A retail investing app is reviewing a feature that surfaces "trending" stocks with a pulsing flame icon, bright green momentum arrows, and a "12,000 buying now" social counter. The team suspects these cues drive buys detached from any judgment about the underlying company. A Salience Normalization Test settles it. Two matched groups of users (or the same users across matched sessions) see the identical set of stocks with identical fundamentals; one group sees the full amplified treatment, the other a normalized presentation — same information, calibrated visual weight, no flame, no live counter.
The result is a measured gap: buy rates on the flagged stocks are markedly higher under amplification, while buy rates on stocks with genuinely strong fundamentals barely move between conditions (illustrative). That pattern is the finding — the decision on the flagged names depends on salience, not substance. The test does not fix anything; it hands the guardrail designers a quantified case for capping or contextualizing those specific cues, and a baseline to re-test against after the fix.
How it works¶
- Fix the substance, vary the salience. Construct two presentations of the identical choice — same options, same underlying information — differing only in cue amplification versus a normalized baseline.
- Define the normalized baseline. The "normal" arm is set to the calibrated range: the presentation whose salience matches the underlying value, so the contrast isolates engineered excess.
- Measure the decision shift. Compare choices, not impressions, across arms; the effect size on decisions is the cue-dependence estimate.
- Read the disproportion signal. A large shift where substantive value is unchanged is the marker of disproportionate, salience-driven response — the output handed to guardrail design.
Tuning parameters¶
- Normalization target — how far the "normal" arm is dialed down. Setting it to a defensible calibration range makes the contrast meaningful; an arbitrary baseline muddies the reading.
- Decision measure — what behavior counts as the outcome (purchase, click-through, hold-vs-sell). Choosing a consequential decision makes the test informative; a trivial proxy weakens it.
- Design — between-subjects (matched groups) versus within-subjects (same users, counterbalanced). Within-subjects is more sensitive but risks carryover; between-subjects avoids that but needs matching.
- Effect threshold — how large a shift counts as cue-dependence worth acting on. A low threshold flags subtle effects but risks noise; a high one only catches gross hijacks.
When it helps, and when it misleads¶
Its strength is evidentiary: it turns an argument about manipulation into a number, and it operationalizes the archetype's core invariant — proportionality — by directly testing whether response tracks value or salience. It is the empirical descendant of the supernormal stimulus demonstration, where a response is shown to follow the exaggerated cue rather than the real referent when the two are pulled apart.[n1] Because it produces a baseline, it also lets a team verify that a later guardrail actually reduced cue-dependence.
Its failure mode is that a test is only as good as its normalized arm and its decision measure: a poorly-chosen baseline, a demand-effect-laden setup, or a trivial outcome can manufacture or hide an effect. It is diagnostic only — it measures the problem and fixes nothing, so a team that runs it and stops has learned without acting. And a single test is a snapshot of one design in one population. The guarding discipline is to anchor the normalized arm in a real calibration range, measure a consequential decision, and treat the result as an input to guardrail design and re-testing rather than a conclusion.
How it implements the components¶
cue_intensity_metric— the amplified and normalized arms are defined by differing measured cue intensity, making intensity the manipulated variable.natural_calibration_range_reference— the normalized arm is set to the calibrated baseline, so the test isolates engineered excess against a defensible reference.harm_and_compulsion_signal— a large decision shift with substance held constant is the disproportionate-response signal the test is built to detect.
It does NOT implement response_ceiling or deamplification_rule — imposing the caps and damping that fix a confirmed hijack is Cue Intensity Cap Protocol's job; this test measures cue-dependence, it does not correct it. It also differs from the inspection-based Supernormal Cue Audit: the audit inventories cues against a range, this runs a behavioral experiment.
Related¶
- Instantiates: Supernormal Cue Guardrail Design — supplies the empirical test that distinguishes salience-driven response from value-driven response.
- Sibling mechanisms: Supernormal Cue Audit · Cue Intensity Cap Protocol · High-Arousal Content Throttle · Cue-Hijack Red-Team Review
Editorial Notes¶
Form Classification¶
Form family: Experiment, Test & Rehearsal
Rationale: Salience Normalization Test operates as an active test, trial, simulation, drill, or rehearsal that generates evidence through a deliberate attempt or perturbation because it compares decisions made under amplified versus normalized cue presentation to measure whether a choice depends on exaggerated salience rather than substantive value.
Independent corroboration: The frozen evidence defines Salience Normalization Test as 'Compares decisions made under amplified versus normalized cue presentation to measure whether a choice depends on exaggerated salience rather than substantive value', so its operative form is Experiment, Test & Rehearsal.
Review outcome: Independent reviewer agreement; high confidence.
Origin Attribution¶
Primary origin: Psychology
Origin pattern: Cross-disciplinary synthesis
Present-day reach: Universal
Rationale: Manipulating cue prominence to measure salience-driven choice is experimental psychology.
Related originating lineages:
- Behavioral Economics — Attention-weighted decision theory materially motivates substantive-value comparison.
- Cognitive Science — Cognitive-science research on representation, learning, and recall supplies a parallel or contributing lineage for the mechanism's defining operation: compares decisions made under amplified versus normalized cue presentation to measure whether a choice depends on exaggerated salience rather than substantive value.
- Statistics & Experimental Design — Controlled presentation contrasts supply causal inference.
Review resolution: Both blind reviewers agree that psychology is the primary historical origin. Explicit reconciliation of alternate_origin_disagreement, origin_mode_disagreement, domain_reach_disagreement starts from reviewer_a's mechanism-specific evidence: Manipulating cue prominence to measure salience-driven choice is experimental psychology. Reviewer A proposed alternates=behavioral_economics, statistics_experimental_design, origin_mode=cross_disciplinary_synthesis, domain_reach=universal, and encyclopedia_synthesis=true; reviewer B proposed alternates=cognitive_science, origin_mode=single_lineage, domain_reach=multi_domain, and encyclopedia_synthesis=true. The final record retains every independently supported alternate from either review (behavioral_economics, statistics_experimental_design, cognitive_science) without an arbitrary cap, selects origin_mode=cross_disciplinary_synthesis to represent the combined lineage evidence, and records domain_reach=universal and encyclopedia_synthesis=true. Present-day transfer is recorded as reach and is not treated as proof of historical origin.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
[n1] The supernormal stimulus — Nobel-laureate ethologist Niko Tinbergen showed that animals often respond more strongly to an exaggerated artificial cue than to the real signal it imitates. A salience normalization test is the applied inverse: pull the exaggerated cue apart from the real value and see which one the response follows. ↩