{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","cell_id":"anchoring_reset__computer_science","arm":"BREADTH_PROBE_ONE_SHOT","candidate_id":"anchoring_reset__computer_science__P1","proposal_index":1,"version":0,"title":"Pre-Score Vulnerability Triage Reset","problem":"Software vulnerability triagers see an automated scanner's severity score before examining the affected service, exploit path, compensating controls, and exposure. The scanner score becomes the starting coordinate for judgment, so later contextual adjustments may remain tied to it even when system-specific evidence points elsewhere.","actors":["Application-security triager","Affected service owner","Automated vulnerability scanner","Release manager responsible for remediation scheduling","Security triage lead"],"observable_state":"For newly detected vulnerabilities, triage notes begin from the scanner score; final priority labels and remediation deadlines cluster near its severity category; independent exploitability estimates are absent; and exceptions are expressed as small adjustments to the score rather than as estimates derived from service context. Tickets, dashboards, and release plans already copy the resulting priority and deadline.","consequence":"A scanner-derived reference can govern remediation order and release work beyond its evidential role, causing teams to defer findings with credible exploit paths or interrupt planned work for findings whose practical reach is constrained.","affected_objective":"Calibrate vulnerability-remediation priority to independently assessed exploitability, impact, exposure, and uncertainty while preserving an auditable decision trail.","intervention":"Insert an anchor-exposure boundary into vulnerability intake. Before revealing the scanner score, a triager records an independent exploit-path assessment, impact range, confidence level, and provisional priority using only the vulnerability description, affected architecture, reachable interfaces, deployed controls, and asset criticality. The score is then revealed and placed beside references from comparable prior findings and available runtime evidence. The triager documents a recalibrated priority and uncertainty band, explains any movement, and traces the result into copied remediation deadlines, dashboards, and release-plan assumptions. High-disagreement or low-confidence cases go to a second reviewer; the reset supplies evidence to the authorized prioritizer but does not make the production decision.","structural_mapping":[{"archetype_element":"Initial reference","domain_realization":"The scanner's displayed severity score and category."},{"archetype_element":"Anchor visibility","domain_realization":"The triage record identifies the scanner, score, display time, scoring inputs, and downstream fields populated from it."},{"archetype_element":"Independent judgment window","domain_realization":"A triager records exploitability, impact, confidence, and provisional priority before the scanner score is displayed."},{"archetype_element":"Alternative Reference Set","domain_realization":"Comparable vulnerabilities in similarly exposed services, observed attack-path evidence, asset criticality, and compensating-control status."},{"archetype_element":"Calibration Evidence","domain_realization":"Reachability checks, deployment configuration, privilege boundaries, exploit prerequisites, telemetry, and documented control effectiveness."},{"archetype_element":"Recalibrated Reference","domain_realization":"An evidence-linked remediation-priority category with an uncertainty band and rationale."},{"archetype_element":"Downstream Anchor Trace","domain_realization":"Review of ticket priority, remediation SLA, dashboard status, exception record, and release-plan entries inherited from the original score."},{"archetype_element":"Escalation Trigger","domain_realization":"Material disagreement between the blind assessment and scanner category, low confidence, contested control effectiveness, or a priority that would trigger an emergency release."}],"mechanism_mapping":[{"mechanism_slug":"blind_independent_estimates","role":"Captures a service-context assessment before the scanner score can pull the triager's judgment toward its category.","counterfactual_removal":"If the score is visible during the first assessment, the supposed independent estimate can remain an adjustment from the anchor."},{"mechanism_slug":"multiple_anchor_comparison","role":"Makes the scanner score one reference among architectural evidence, comparable cases, and runtime observations.","counterfactual_removal":"Without alternative references, reviewers can identify the anchor yet still lack another coordinate from which to recalibrate."},{"mechanism_slug":"downstream_anchor_trace","role":"Finds and revises deadlines, dashboard labels, and release assumptions that copied the scanner-derived priority.","counterfactual_removal":"Without tracing, the old priority can continue governing work even after the official assessment changes."}],"causal_chain":["The scanner score appears before contextual investigation.","The score defines the initial severity coordinate and perceived adjustment range.","The triager's contextual judgment remains partly organized around that coordinate.","The anchored priority propagates into remediation deadlines, dashboards, and release planning.","A pre-score assessment creates a judgment not derived from the scanner reference.","Comparison with multiple references and system evidence exposes agreement, disagreement, and uncertainty.","An authorized reviewer adopts or rejects a documented recalibrated priority.","Dependent workflow artifacts are updated so the original anchor does not remain operational."],"baseline":"The scanner creates a ticket containing its score and severity category; a triager reviews service context with that score visible, edits the category or deadline when warranted, and downstream systems copy the final ticket fields. The baseline records the outcome but does not preserve a pre-anchor estimate or distinguish evidence-derived judgment from adjustment around the scanner score.","nearest_rivals":["Contextual modification of the scanner's severity score after it is displayed","A risk-based vulnerability-management model that combines technical severity with asset criticality","A second-reviewer approval requirement for priority changes","A generic anchoring-bias checklist attached to the triage form"],"remaining_contrastive_claim":"The candidate's distinctive operational claim is that exposure order and downstream inheritance are part of the defect: it captures a contextual estimate before score exposure, calibrates rather than merely modifies the score, and repairs artifacts that already inherited the anchored priority. A richer risk formula or second approval performed with the original score visible does not supply that independent reference.","authority_safety":{"decision_authority":"The security triage lead owns the assessment procedure, while the designated vulnerability-risk owner retains authority over final remediation priority; service owners and release managers retain their existing operational authorities.","authorized_first_step":"The security triage lead may run a shadow-mode pilot in which selected findings receive a pre-score assessment and comparison record, without changing production tickets, deadlines, access controls, scanner configuration, or release schedules.","excluded_actions":["Automatically lowering or raising production remediation priority","Suppressing scanner findings or hiding scores from personnel who require them for current duties","Changing remediation deadlines or release schedules without the existing risk owner's approval","Treating the independent estimate as authoritative solely because it was blind","Disclosing sensitive architecture or vulnerability details outside approved security channels"],"halt_rollback":"Stop the pilot if score masking obstructs an active incident, if required evidence cannot be provided within existing access rules, if triagers can infer the score before recording their estimate, or if shadow records leak into production workflow. Restore the ordinary score-visible form and discard pilot-only workflow fields under the applicable retention rules; production priorities remain unchanged."},"negative_tests":{"strongest_counterevidence":"Pre-score and score-visible assessments may show no systematic shift attributable to score exposure, while disagreements may instead track missing architectural evidence or differing risk policies.","problem_falsifier":"The anchoring problem is falsified for the tested workflow if triagers already form and timestamp independent contextual assessments before viewing the scanner score, final judgments track admissible system evidence rather than proximity to the score, and downstream artifacts do not inherit the score before review.","intervention_falsifier":"The intervention is falsified if blinded assessments are reproducible and informed but revealing the score does not measurably change priority, confidence, or rationale relative to the existing workflow, or if any changes are no better aligned with the predefined evidence rubric and later authorized adjudication.","risks":["Masking the score may omit useful technical information encoded in its components, producing a weaker first assessment.","Triagers may infer the hidden category from wording or workflow cues, creating false confidence in blindness.","A counter-anchor may form around asset criticality or a salient prior incident.","The extra assessment step may delay urgent handling.","Reviewers may overcorrect away from a valid scanner score merely to demonstrate independence.","Comparable cases may be selectively chosen or may not match the current architecture.","Shadow records may expose sensitive vulnerability or topology information.","An uncertainty band may be collapsed into a crisp deadline by downstream tooling."]},"next_evidence_step":"Run a shadow-mode test on the next 12 eligible non-incident findings or for two weeks, whichever occurs first. Randomly assign each finding to either the ordinary score-visible first assessment or the pre-score form, then have both groups complete the same evidence rubric and an authorized adjudicator review masked rationales. Record category movement after score reveal, confidence, rationale completeness, time burden, agreement with adjudication, and whether downstream fields would differ. Do not alter production priorities. Proceed beyond shadow mode only if blindness is credible, evidence capture remains compliant, urgent handling is not delayed, and results distinguish anchor carryover from missing information or policy disagreement.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Not assessed against other proposals under runtime isolation; this candidate realizes the archetype specifically as an exposure-order and inheritance-control intervention for automated vulnerability-severity scores in software remediation workflows.","revision_record":{"parent_version":null,"progress_targets_addressed":[],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}