{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"negative_space_design__linguistics_semiotics:P4:v0","cell_id":"negative_space_design__linguistics_semiotics","search_queries":["annotation bias pre-annotation labels anchoring corpus annotation study","blind annotation existing labels bias quality control corpus linguistics","INCEpTION curation annotation workflow adjudication documentation","WebAnno annotation curation workflow disagreement","\"Anchoring and Agreement in Syntactic Annotations\"","\"Quality and Efficiency of Manual Annotation\" \"Pre-annotation Bias\"","annotation tool blind mode hide preannotations review reveal","NLP annotation blind annotation adjudication existing labels","\"Evaluating the impact of pre-annotation\" annotation speed potential bias full text","site:w3.org/WAI/WCAG22 focus order status messages","site:aclanthology.org \"Analyzing Dataset Annotation Quality Management in the Wild\"","annotation experiment pre-annotation bias dependency syntax manual annotation ACL","doi \"Evaluating the impact of pre-annotation on annotation speed\" Lingren","site:academic.oup.com/jamia \"pre-annotation\" Lingren","\"Evaluating the impact of pre-annotation\" Lingren Deleger"],"sources":[{"source_id":"S1","title":"Anchoring and Agreement in Syntactic Annotations","publisher":"Association for Computational Linguistics","url":"https://aclanthology.org/D16-1239/","source_class":"PRIMARY_RESEARCH","publication_date":"2016-11","accessed_at":"2026-08-03","claims_supported":["Editing parser output can anchor syntactic annotators toward pre-existing analyses.","The study reported lower-quality annotations and overestimated parser performance relative to human-based annotation.","Independent human annotation is an established comparator for syntactic annotation workflows."]},{"source_id":"S2","title":"Quality and Efficiency of Manual Annotation: Pre-annotation Bias","publisher":"European Language Resources Association","url":"https://aclanthology.org/2022.lrec-1.312/","source_class":"PRIMARY_RESEARCH","publication_date":"2022-06","accessed_at":"2026-08-03","claims_supported":["A four-annotator dependency-syntax experiment directly compared pre-parsed and from-scratch annotation.","Pre-annotation influenced judgments and raised agreement, but did not reduce measured accuracy in that setting.","From-scratch annotation took about 1.7 times as long, establishing a material workload tradeoff.","The PDT-C team explicitly sought the best workflow for a two-million-token annotation project, evidencing an identifiable user and authorizer need."]},{"source_id":"S3","title":"Evaluating the impact of pre-annotation on annotation speed and potential bias: natural language processing gold standard development for clinical named entity recognition in clinical trial announcements","publisher":"Journal of the American Medical Informatics Association","url":"https://pubmed.ncbi.nlm.nih.gov/24001514/","source_class":"PRIMARY_RESEARCH","publication_date":"2013-09-03","accessed_at":"2026-08-03","claims_supported":["A double-annotation experiment treated unlabeled text as the comparator to pre-annotated text.","Dictionary pre-annotation saved 13.85% to 21.5% per entity without a statistically significant loss in agreement or annotator performance in that clinical NER setting.","The effect of withholding prior answers is task-dependent rather than uniformly beneficial."]},{"source_id":"S4","title":"INCEpTION User Guide, version 41.2","publisher":"The INCEpTION Team","url":"https://inception-project.github.io/releases/41.2/docs/user-guide.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-07-21","accessed_at":"2026-08-03","claims_supported":["INCEpTION already separates annotator, curator, and manager roles.","Its standard workflow collects annotations before curation, then displays agreement and disagreement for a curator's final selection.","Project managers and curators are technically credible authorizers for a reversible annotation-workflow pilot."]},{"source_id":"S5","title":"WebAnno User Guide, version 3.6.5","publisher":"The WebAnno Team","url":"https://webanno.github.io/webanno/releases/3.6.5/docs/user-guide.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2020-04-05","accessed_at":"2026-08-03","claims_supported":["WebAnno implements independent annotation followed by a curation interface showing annotators' analyses in read-only panes.","Agreement can be auto-merged while disagreements remain for curator resolution.","Anonymous curation and role-based workflow are existing adjacent bias-control practices."]},{"source_id":"S6","title":"Analyzing Dataset Annotation Quality Management in the Wild","publisher":"MIT Press","url":"https://aclanthology.org/2024.cl-3.1/","source_class":"PRIMARY_RESEARCH","publication_date":"2024-09","accessed_at":"2026-08-03","claims_supported":["A review of 591 text-dataset publications found that annotation quality management commonly involves agreement, adjudication, and validation.","Thirty percent of reviewed studies were assessed as having subpar quality-management effort.","The paper identifies recurring mistakes in agreement and annotation-error-rate analysis, supporting careful outcome design."]},{"source_id":"S7","title":"Web Content Accessibility Guidelines (WCAG) 2.2","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/WCAG22/","source_class":"STANDARD","publication_date":"2024-12-12","accessed_at":"2026-08-03","claims_supported":["Sequential keyboard focus must preserve meaning and operability.","Headings and labels must describe purpose, and status messages must be programmatically determinable.","The staged reveal can be implemented accessibly, but compliance requires explicit focus, labeling, keyboard, and status-message testing."]},{"source_id":"S8","title":"Adjudication Mode","publisher":"Potato Annotation","url":"https://potatoannotator.readthedocs.io/en/stable/administration/adjudication/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Potato implements separate annotation and adjudication stages with designated adjudicator accounts.","It tracks provenance, agreement, timing, confidence, override notes, and disagreement categories.","It can hide annotator identities to reduce bias, demonstrating configurable withholding of authority cues as existing product practice."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible in controlled syntactic-annotation research: S1 found anchoring and quality distortion when annotators edited parser output. S2 independently confirms that pre-annotation influences judgments. However, S2 and S3 found no measured accuracy penalty and substantial time savings in their settings, so prevalence and consequence for targeted audits—not existence of the mechanism—remain uncertain.","source_ids":["S1","S2","S3","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"PDT-C researchers explicitly needed to select a workflow for a large dependency corpus, while INCEpTION identifies managers and curators with workflow and final-label authority. Current tools also document provenance and adjudication needs. No source records a corpus team requesting this exact blank-first/reveal-later feature or committing resources to it.","source_ids":["S2","S4","S6","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"From-scratch independent annotation followed by adjudication","similarity":"Already creates answer-free judgments, preserves multiple annotator outputs, and exposes disagreements for later resolution.","remaining_difference":"The proposal applies that logic only to bounded disputed decisions inside an audit and reveals inherited labels immediately after the first-pass commitment rather than requiring full independent reannotation.","source_ids":["S2","S3","S4","S5","S8"]},{"name":"Parser-output editing versus human-based syntactic annotation","similarity":"Directly tests whether visible prior analyses anchor dependency judgments and uses unexposed human annotation as the comparator.","remaining_difference":"It is an experimental annotation-production comparison, not a productized two-stage audit bay with context requests, abstention, immutable provisional logs, and post-commit reveal.","source_ids":["S1"]},{"name":"INCEpTION/WebAnno/Potato curation workflows","similarity":"These products already separate roles and phases, preserve provenance, display disagreement, support final adjudication, and sometimes suppress identity cues.","remaining_difference":"Their documented curator view reveals annotator answers during adjudication; no relied-upon source documents requiring an adjudicator's own first-pass analysis before those answers appear.","source_ids":["S4","S5","S8"]},{"name":"Reveal-first pre-annotation with quality checks","similarity":"Uses existing analyses and automated support to improve speed, consistency, and sometimes quality.","remaining_difference":"It intentionally exposes prior answers from the start and therefore cannot independently attribute an initial reviewer judgment.","source_ids":["S2","S3"]}],"distinctive_claim_remaining":"For bounded morphosyntactic audits, requiring a time-stamped provisional analysis, context-insufficiency decision, or abstention before revealing inherited labels will reduce adoption of plausible inherited errors and surface more independently generated disagreements than a reveal-first interface with an instruction to judge independently, without materially reducing reference agreement or creating unacceptable time, context-repair, abstention, or accessibility costs.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Existing annotation systems demonstrate role separation, independent annotation, curation queues, provenance, timing, confidence, disagreement categorization, and configurable identity withholding. A read-only staged prototype is technically conventional. Missing evidence includes an implementation of the exact commit-then-reveal interaction, integration complexity for dependency-tree editors, usable context-selection rules, security review, assistive-technology performance, and behavior under prior corpus familiarity.","source_ids":["S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Avoiding inherited syntactic errors could improve corpus validity, but externally observed effects conflict across tasks and realized downstream impact is unmeasured.","source_ids":["S1","S2","S3","S6"]},"stakeholder_pull":{"score":3,"rationale":"Large corpus projects and documented manager/curator roles establish a real workflow owner and quality need, but no adopter has requested or funded this exact intervention.","source_ids":["S2","S4","S8"]},"incremental_advantage":{"score":3,"rationale":"The intervention may preserve most of reveal-first curation's recoverability while improving first-pass provenance, but its advantage over an instruction-only interface and ordinary independent annotation is untested.","source_ids":["S1","S2","S4","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"Independent annotation followed by visible adjudication is established practice; the main remaining distinction is a narrower, staged audit interaction rather than a new quality-control principle.","source_ids":["S2","S4","S5","S8"]},"technical_implementability":{"score":4,"rationale":"Current open annotation systems already possess most required primitives; the new work is principally state gating, immutable logging, reveal controls, dependency-specific UI, and accessibility validation.","source_ids":["S4","S5","S7","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"Managers and curators are identifiable authorizers and a read-only pilot avoids production-label authority, but no partner, data-license review, institutional approval, or human-subjects determination is documented.","source_ids":["S4","S8"]},"evidence_readiness":{"score":3,"rationale":"The claim has clear comparators, measurable outcomes, and directly relevant prior experiments, but determining incremental value requires live reviewer behavior rather than further web research.","source_ids":["S1","S2","S3","S6"]},"safety_net_benefit":{"score":2,"rationale":"Read-only operation, abstention, context requests, and separate adjudication limit harm, but the intervention is a corpus-quality control rather than a direct safety system and can omit necessary context or burden assistive-technology users.","source_ids":["S7","S8"]},"scalability":{"score":4,"rationale":"Once implemented, state gating and audit logging can be reused across many items and projects, although extra independent judgment time may constrain broad deployment.","source_ids":["S2","S4","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Build a read-only prototype and run the proposed 48-item, four-reviewer controlled study, including item preparation, accessibility checks, analysis, and adjudicator review.","confidence":"MODERATE","assumptions":["Approximately 120-220 hours of engineering, experiment design, item preparation, analysis, and project management.","Four qualified reviewers complete 192 total decisions plus training and debriefing.","Existing adjudicated data and an annotation-platform test instance are available without acquisition fees.","The estimate is resource-equivalent, not a vendor quotation."],"source_ids":["S2","S4","S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production-grade integration into one annotation platform with access controls, immutable provenance, dependency-tree interaction, context-request logic, rollback, security review, accessibility remediation, and administrator documentation.","confidence":"LOW","assumptions":["One existing platform and one corpus schema are in scope.","No rewrite of the underlying annotation editor is required.","Institutional hosting, identity management, and source-data permissions already exist.","Security and accessibility testing identify remediable rather than architectural failures."],"source_ids":["S4","S5","S7","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Launch on one corpus program, including workflow configuration, reviewer training, an initial bounded audit batch, adjudication capacity, monitoring, and evaluation.","confidence":"LOW","assumptions":["Roughly 500-2,000 targeted decisions rather than corpus-wide reannotation.","A corpus manager and adjudication lead contribute part-time.","Blank-first adds review time relative to reveal-first, consistent with from-scratch annotation's measured time penalty.","No sensitive-data relocation or new licensing is required."],"source_ids":["S2","S4","S8"]},"annual_recurring":{"band_2026_usd":"10K_TO_50K","scope":"Software maintenance, accessibility regression testing, audit-log retention, administrator support, refresher training, and a limited annual targeted-review program.","confidence":"LOW","assumptions":["One maintained deployment and periodic audits, not continuous high-volume annotation.","Reviewer labor scales with the number and complexity of audited decisions.","Hosting and core annotation-platform operations are already funded.","No dedicated full-time engineering or adjudication staff is included."],"source_ids":["S2","S4","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Controlled syntactic research directly demonstrates anchoring from visible parser analyses, while later studies confirm judgment influence even where accuracy loss was absent.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Dependency-corpus programs have explicitly evaluated annotation workflow choices, and current platforms assign project managers and curators authority over workflow and final curation.","source_ids":["S2","S4","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal specifies a staged commit-before-reveal mechanism and can be compared against reveal-first plus an independence instruction using error adoption, reference agreement, surfaced disagreement, time, repair, abstention, and usability outcomes.","source_ids":["S1","S2","S3"]},"bounded_next_evidence_step":{"status":"YES","reason":"A 48-item read-only, counterbalanced study with four reviewers, controlled inherited-label perturbations, two interface conditions, fixed outcomes, and explicit stop thresholds is bounded.","source_ids":["S1","S2","S3","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"Read-only data and separate adjudication bound corpus risk, and WCAG supplies implementation criteria, but actual corpus permissions, restricted-data exposure, institutional human-subjects status, and assistive-technology performance require partner-specific review.","source_ids":["S4","S7","S8"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The bands have explicit labor and scope assumptions and research documents a meaningful time penalty for from-scratch work, but no direct 2026 wage, integration, or institutional-cost evidence was found.","source_ids":["S2","S4","S7","S8"]}},"next_evidence_step":"With a corpus partner, obtain data-use and human-subjects determinations, then copy 48 previously adjudicated dependency decisions into a non-production test set spanning clear, ambiguous, and context-dependent cases. For a predeclared subset, replace the displayed inherited label only in the test copy with a plausible known-wrong alternative while retaining the reference and provenance. Four eligible reviewers each assess every item once; matched blocks counterbalance (A) blank-first commit-then-reveal and (B) reveal-first with an explicit instruction to judge independently, giving two reviews per item per condition. Preserve identical sentences, guidelines, context-request options, abstention options, and adjudication criteria. Primary outcomes are adoption of planted inherited errors and agreement of the initial response with the independent reference. Secondary outcomes are independently surfaced disagreements, context requests, insufficiency/abstention, time, post-reveal revision, shallow reveal-seeking responses, focus/keyboard/screen-reader failures, and reviewer comprehension of the withholding boundary. Falsify the incremental claim if blank-first does not materially reduce planted-error adoption or improve independently attributable error detection; also reject or redesign if reference agreement falls by more than 5 percentage points, median time rises by more than 50%, context-repair or abstention rises by more than 15 percentage points, or any unresolved accessibility or necessary-context failure occurs. Do not promote any pilot response into the corpus.","blocking_evidence":["No corpus organization has committed to adopt, authorize, or fund the pilot.","The exact staged audit workflow was not found as an evaluated product feature, but bounded search cannot establish world novelty.","Incremental benefit over reveal-first plus an independence instruction requires live testing.","The frequency and consequence of answer exposure specifically in targeted morphosyntactic audits are not externally quantified.","The correct load-bearing context policy for complex or cross-sentence analyses is unvalidated.","Corpus licensing, confidentiality, security, and institutional human-subjects requirements are partner-specific and unresolved.","Accessibility of dynamic withholding, commit, reveal, and focus transitions has not been tested with users or assistive technologies.","The resource bands lack direct 2026 labor-rate or vendor-cost validation."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation establishes only that anchoring research, from-scratch independent annotation, staged annotation-to-curation workflows, provenance-aware adjudication, and selective identity hiding are close prior art. It did not measure world novelty, patentability, freedom to operate, market size, realized impact, or exhaustive product/repository coverage. No relied-upon source documented the exact targeted commit-before-reveal audit interaction, but absence from this bounded search is not evidence of novelty.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Secure a named corpus partner, corpus-lead authorization, data-use clearance, and any required human-subjects determination.","Pre-register the two-condition, controlled-error pilot, primary estimand, analysis method, and falsification thresholds.","Implement and accessibility-test immutable first-pass logging, context requests, abstention, reveal recovery, keyboard focus, screen-reader announcements, and rollback.","Run the bounded live-reviewer experiment and report planted-error adoption, reference agreement, surfaced disagreements, time, repair, abstention, and usability by condition.","Demonstrate incremental value over both reveal-first plus instruction and the established alternative of ordinary independent annotation followed by curation.","Replace resource-equivalent estimates with tracked engineering, reviewer, adjudication, compliance, and maintenance hours from the pilot."],"reason":"Web evidence verifies a real but context-dependent anchoring mechanism and shows that most workflow primitives already exist. It also creates a substantial prior-art collision with independent annotation followed by curation and includes counterevidence that pre-annotation can save time without reducing measured quality. The remaining proposition is behavioral and partner-specific; it cannot be resolved by additional bounded web research and requires live reviewer testing plus authority and accessibility checks."},"proposal_index":4}