{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__criminology_forensic","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"brg_forensic_error_discovery_league_p1_v0","proposal_index":1,"version":0,"title":"Bounded Forensic Error-Discovery League","problem":"A regional forensic consortium allocates scarce method-validation time, training places, and instrument-upgrade credits partly from laboratory performance records. When turnaround and an appearance of error-free operation influence standing, participating units can improve their position by avoiding difficult reviews, classifying discrepancies narrowly, or withholding lessons from other units. The resulting rivalry rewards a clean-looking record rather than demonstrated ability to find, characterize, and safely communicate analytical or chain-of-custody weaknesses.","actors":["Forensic laboratory analytical teams","Laboratory quality managers","Regional forensic-science oversight board","Independent challenge-set designers and scorers","Defense and prosecution representatives serving as safeguard observers","People whose cases or reference materials could be affected by forensic findings","Accreditation and funding officials"],"observable_state":"The consortium can observe upgrade and training resources being allocated among units; performance reports emphasizing turnaround and low reported-error counts; differences between defects found during blinded rechecks and events entered into corrective-action records; reluctance to share failure analyses; and disputes over whether laboratory standing reflects analytical quality or reporting strategy.","consequence":"Units can gain scarce resources and reputation without becoming better at detecting error, while difficult evidence, transferable failure lessons, and calibrated uncertainty receive less attention. If used in casework, these incentives could also shift costs to defendants, victims, investigators, and courts through delayed correction or overstated confidence.","affected_objective":"Make scarce forensic-development resources reward reproducible error detection, calibrated interpretation, and safe corrective learning while preserving case rights, evidence integrity, and future contestability.","intervention":"Create a no-case-impact error-discovery league among qualified forensic teams. Teams receive matched, de-identified challenge packets containing known-source samples, synthetic analytical outputs, and seeded documentation or chain-of-custody defects. A frozen rulebook defines eligibility, equal analysis-hour limits, permitted tools, prohibited contact or data access, scoring, appeals, and incident handling. Scores combine correctly characterized defects, calibrated confidence, reproducibility on a held-out packet, appropriate escalation, and penalties for false accusations, unsupported conclusions, leakage, or evidence alteration. Independent scorers audit leading submissions. Scarce development credits are divided among complementary achievements rather than assigned to one overall winner. Any incidental indication concerning a real case is sealed and referred outside the contest; it earns no points and cannot be acted upon by contest officials. After the round, the oversight board reviews gaming, burdens, spillovers, and whether the scoring rule should be revised or the league retired.","structural_mapping":[{"archetype_element":"Rivalry purpose statement","domain_realization":"Use rivalry to reveal forensic-process weaknesses and disciplined error-characterization practices, not to minimize reported errors or maximize accusation counts."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"A fixed pool of method-validation time, specialist training places, and instrument-upgrade credits."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Accredited units or supervised teams that accept conflict disclosure, confidentiality, equal-resource, and reproducibility requirements."},{"archetype_element":"Contest arena boundary","domain_realization":"Analysis is limited to copied challenge materials; contacting case actors, entering operational evidence systems, identifying personnel, altering source records, and using contest findings in litigation or discipline are prohibited."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Verified defect detection is balanced against false-positive findings, confidence calibration, reproducibility, documentation quality, and compliance with safe escalation procedures."},{"archetype_element":"Fair process and due process layer","domain_realization":"Rules and challenge-construction procedures are frozen before entry; teams receive scored evidence and may appeal scoring errors to a reviewer separate from designers and scorers."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Packet watermarking, access logs, fixed cross-team communication rules, anomaly screening, and graduated penalties address answer sharing, leakage, tampering, and interference with rivals."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Only synthetic, known-source, or safely copied materials affect scoring; identities are masked, real-case signals are quarantined, and contest results cannot determine individual case or employment outcomes."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"Every team receives the same analysis-hour ceiling, packet access period, and approved resource envelope."},{"archetype_element":"Winner power and recalibration","domain_realization":"Awards are distributed across complementary achievements, scoring data do not confer rulemaking authority, and an independent post-round review can revise or terminate the arena."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Freezes eligibility, allowed actions, scoring, tie-breaks, confidentiality, incident handling, and appeal rights before teams see the packets.","counterfactual_removal":"Without it, organizers could reinterpret errors or change scoring after seeing which laboratory benefits, and teams would lack a contestable arena boundary."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Makes comparative performance observable while requiring independent reproduction and held-out verification before any award.","counterfactual_removal":"Without audit, aggressive overcalling, fabricated documentation, or challenge-specific gaming could produce the leading score."},{"mechanism_slug":"multiple_award_or_portfolio_selection","role":"Splits development credits across complementary strengths such as analytical detection, documentation control, and calibrated interpretation.","counterfactual_removal":"A single aggregate winner could dominate future resources and induce teams to neglect capabilities poorly represented by the headline score."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Equalizes analysis hours and approved resource access so the contest tests practice within a bounded envelope rather than laboratory wealth.","counterfactual_removal":"Better-resourced units could win through overtime, personnel volume, or proprietary access, provoking an expenditure race unrelated to error-discovery quality."},{"mechanism_slug":"anti_collusion_monitoring","role":"Uses packet-access and submission patterns to flag possible answer sharing or coordinated score management for separate investigation.","counterfactual_removal":"Teams could appear independently successful while sharing discoveries or arranging complementary underperformance; the ranking would cease to represent rivalry."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Publishes graduated consequences for leakage, tampering, impersonation, interference, unsupported identification, and prohibited access.","counterfactual_removal":"Off-arena conduct could become a viable winning strategy, while ad hoc enforcement would invite selective punishment."},{"mechanism_slug":"post_contest_impact_review","role":"Tests whether the contest rewarded the stated forensic objective, generated harmful pressure, disadvantaged certain units, or created resource lock-in, then revises the next round.","counterfactual_removal":"A gamed or harmful scoring system could persist merely because it produced a clear ranking."}],"causal_chain":["Scarce development resources and institutional standing make forensic units strategic rivals.","The existing emphasis on speed and clean-looking records permits advantage through narrow error classification, avoidance, or silence.","The league redirects a bounded portion of that rivalry toward finding and responsibly characterizing controlled defects.","Hidden challenge construction and balanced scoring make indiscriminate error claims costly and verified, reproducible findings valuable.","Equal resource limits reduce the advantage of bankroll and dampen an analytical spending race.","Audits, access-pattern screens, appeal rights, and foul penalties constrain gaming, collusion, sabotage, and organizer discretion.","Portfolio awards preserve multiple capabilities and prevent one score from conferring control over the field.","Post-round review compares realized conduct with the stated purpose and revises or retires the contest if winning separates from safe forensic contribution."],"baseline":"Continue routine internal quality assurance, periodic proficiency exercises, accreditation review, and discretionary allocation of development resources, with laboratory standing influenced by throughput and reported corrective events. This baseline supports compliance but does not deliberately make validated discovery and sharing of weaknesses the route to a scarce reward.","nearest_rivals":["A noncompetitive independent forensic audit, which can inspect quality without inducing rivalry but may sample fewer approaches and provides no comparative incentive for teams to surface weaknesses.","Ordinary proficiency testing, which evaluates whether analysts reach expected answers but need not reward discovery of hidden procedural defects, calibrated escalation, or transferable failure analysis.","A conventional laboratory leaderboard based on turnaround, backlog, or low reported-error counts, which is simpler but preserves incentives to optimize visible proxies and suppress adverse information.","Direct discretionary grants for quality improvement, which avoid contest pressure but depend on reviewers predicting which teams and methods will produce useful learning."],"remaining_contrastive_claim":"The proposal's distinctive testable claim is not that competition improves forensic quality generally. It is that, where scarce development resources already create strategic comparison, a governed contest can make verified weakness discovery and safe disclosure a more reliable route to those resources than maintaining an apparently clean record, provided scoring, resource, case-rights, and recalibration boundaries hold.","authority_safety":{"decision_authority":"A regional forensic-science oversight board may authorize only the controlled challenge protocol and allocation of pilot participation resources; independent legal or case-review authorities retain all power over real evidence, disclosure duties, litigation, accreditation, employment, and discipline.","authorized_first_step":"Approve a preregistered, no-stakes dry run using synthetic and known-source packets, volunteer teams, an independent scorer, equal time limits, and a written quarantine procedure for incidental real-case signals.","excluded_actions":["Opening, reopening, charging, dismissing, or otherwise changing a criminal case","Using contest scores as evidence in court, accreditation, promotion, discipline, or termination","Naming analysts, laboratories, defendants, victims, or investigators in public standings","Accessing or altering operational evidence, laboratory, police, or court systems","Awarding actual upgrade funds before the dry run and safeguard review","Treating anomaly screens as proof of collusion or misconduct","Allowing competitors, scorers, or challenge designers to investigate incidental real-case signals"],"halt_rollback":"Immediately freeze scoring and packet access upon identity leakage, operational-system access, evidence-integrity concern, unmanaged case-specific information, credible scorer conflict, or unequal resource exposure. Preserve an access log, notify the independent safeguard reviewer, quarantine affected materials, and void the round if blinding or comparability cannot be restored. Because the first step is no-stakes and uses copies, rollback consists of revoking access, destroying controlled working copies under the protocol, withholding rankings, and making no resource allocation."},"negative_tests":{"strongest_counterevidence":"Teams score better chiefly by overcalling defects, exploiting packet-construction artifacts, sharing answers, or concentrating exceptional personnel in the contest, while held-out reproduction, false-positive control, ordinary quality work, or participant candor deteriorates.","problem_falsifier":"The inferred problem is unsupported if blinded rechecks, corrective-action records, and allocation decisions show that units already gain rather than lose standing from transparent weakness discovery, difficult reviews are not avoided, and scarce development resources do not create strategic comparison among units.","intervention_falsifier":"The intervention fails its mechanism test if, under equal resources and blinded scoring, contest-framed teams do not produce more verified and reproducible defect characterizations than matched non-ranked review teams, or if any gain depends on unacceptable false accusations, leakage, collusion, workload displacement, or circumvention of case-rights safeguards.","risks":["The scoring proxy could reward defect quantity over forensic relevance.","Synthetic packets may not represent the ambiguity and workflow constraints of real examinations.","Public or internal rankings could stigmatize laboratories or analysts despite masking rules.","Resource caps may advantage incumbents that already possess specialized tools or reusable workflows.","Teams could collude, leak packet contents, or infer challenge-construction signatures.","Competitive pressure could divert staff from routine quality work or discourage cooperation outside the arena.","An incidental real-case signal could be mishandled, suppressed, or improperly publicized.","Portfolio awards could become token distribution rather than a genuine down-select."]},"next_evidence_step":"Run one bounded, no-stakes matched-packet study. Randomly assign equivalent synthetic defect packets to volunteer teams under either the proposed governed-contest instructions or a non-ranked independent-review instruction, while holding time, tools, and scorer access constant. Pre-register verified defect characterization, false-positive findings, confidence calibration, held-out reproducibility, hours used, rule violations, cross-team information leakage, and participant-reported pressure. An independent safeguard reviewer examines only protocol compliance and adverse events. Stop before rankings or resource awards if blinding, equal exposure, or quarantine controls fail.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"No earlier proposal is part of this sealed request, so contrast with prior proposals is not applicable; this candidate is defined solely from the supplied archetype, mechanisms, and forensic-domain card.","revision_record":{"parent_version":null,"progress_targets_addressed":["Initial complete proposal","Concrete scarce prize and rival actors","Observable domain state and causal chain","Explicit authority boundary and reversible first step","Contrast with strong non-rival and rival baselines","Problem and intervention falsifiers"],"conceptual_changes":["Initial version; no parent revision."],"operational_changes":["Specified synthetic matched packets, equal analysis-hour limits, independent scoring, appeals, quarantine, and no-stakes rollback."],"evidence_changes":["Defined a bounded matched-packet mechanism test without live cases or operational evidence."],"claim_changes":["Limited the claim to incentive redirection under governed rivalry; made no claim of novelty, prevalence, demand, effect size, or established efficacy."]}}