{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"deadweight_loss_reduction__linguistics_semiotics","trajectory_id":"R","attempt_index":0,"archetype_slug":"deadweight_loss_reduction","domain_slug":"linguistics_semiotics","decision":"CANDIDATE","problem_id":"valid_dialect_variant_rejection_in_language_assessment","causal_lever_id":"variant_equivalence_scoring_rubric_pilot","proposal":{"problem":"HYPOTHESIS: Language assessments used for placement, certification, or access sometimes reject dialectal or regional forms that satisfy the task's semantic and pragmatic requirements solely because they depart from a canonical-form whitelist. This coarse eligibility rule can misclassify communicatively competent candidates while serving the legitimate purposes of consistent scoring and verifying task-relevant proficiency.","actors_substrate":["assessment candidates using standard and nonstandard varieties","test designers and psychometricians","raters and appeals staff","institutions relying on scores","communities whose varieties are evaluated"],"observable_state":"After controlling for task-relevant comprehension and production, otherwise adequate responses containing specified noncanonical variants receive lower scores, trigger retakes or appeals, or change placement disproportionately.","consequence":"Competent candidates lose access or spend time on retesting; institutions discard usable linguistic capability; affected communities bear concentrated conformity costs.","affected_objective":"Accurate, reliable, and equitable measurement of task-relevant communicative proficiency.","structural_mapping":[{"archetype_element":"avoidable wedge","domain_realization":"HYPOTHESIS: A canonical-form-only scoring rule rejects some functionally equivalent variants independently of communicative performance.","claim_kind":"HYPOTHESIS"},{"archetype_element":"blocked beneficial activity","domain_realization":"INFERENCE: Misclassified candidates cannot enter appropriate instruction, certification, or language-dependent roles despite adequate task performance.","claim_kind":"INFERENCE"},{"archetype_element":"protected purpose","domain_realization":"Comparable scoring, intelligibility, construct validity, and dependable certification remain non-negotiable.","claim_kind":"INFERENCE"},{"archetype_element":"wedge-specific redesign","domain_realization":"Admit prevalidated variant classes through an equivalence rubric while retaining performance thresholds and human review for ambiguous cases.","claim_kind":"HYPOTHESIS"},{"archetype_element":"bounded monitoring and reversal","domain_realization":"Pilot on archived responses plus a limited live cohort; compare validity, reliability, incidence, appeals, and downstream performance, then expire or revert automatically.","claim_kind":"HYPOTHESIS"}],"component_map":[{"component":"Distortion Map","status":"direct","domain_realization":"Trace canonical-only rules to variant rejection, score changes, retakes, and denied placement."},{"component":"Protected Constraint Safeguard","status":"direct","domain_realization":"Preserve intelligibility, construct validity, scoring consistency, and minimum proficiency."},{"component":"Surplus Estimate","status":"adapted","domain_realization":"Estimate corrected classifications, avoided retakes, recovered access, and review burden with uncertainty ranges."},{"component":"Affected-Party Incidence Map","status":"direct","domain_realization":"Compare effects across variety groups, candidates, raters, institutions, and downstream interlocutors."},{"component":"Redesign Lever","status":"direct","domain_realization":"Replace the canonical-form whitelist with a task-conditioned variant-equivalence rubric."},{"component":"Distributional Review","status":"direct","domain_realization":"Test whether gains or new errors concentrate by dialect, race, region, class, disability, or language background."},{"component":"Behavioral Response Model","status":"adapted","domain_realization":"Anticipate coaching to the expanded rubric, rater inconsistency, strategic ambiguity, and increased demand."},{"component":"Implementation Boundary","status":"direct","domain_realization":"Limit the pilot to named tasks, variant classes, sites, and one scoring cycle."},{"component":"Monitoring and Rebound Check","status":"direct","domain_realization":"Track reliability, validity, appeals, processing time, subgroup errors, and downstream communication failures."},{"component":"Rollback or Adjustment Rule","status":"direct","domain_realization":"Restore the prior rubric if validity, reliability, or protected-group outcomes cross preset bounds."},{"component":"Cost–Benefit Assessment Frame","status":"adapted","domain_realization":"Compare corrected access and avoided retesting against validation, training, review, and error costs."},{"component":"Price-Wedge Diagnostic","status":"incompatible","domain_realization":"No administered price is posited; fees and retake costs are consequences, not the causal wedge."},{"component":"Friction Source Breakdown","status":"direct","domain_realization":"Separate rubric exclusion from rater calibration, administrative delay, ambiguous prompts, and genuine skill deficits."},{"component":"Compensating Adjustment Plan","status":"adapted","domain_realization":"Offer no-fee rescoring or appeal for pilot errors and accessible guidance for candidates."},{"component":"Legitimacy and Authority Review","status":"direct","domain_realization":"Confirm the assessment owner may alter scoring without exceeding certification or accreditation mandates."},{"component":"Sensitivity Analysis","status":"direct","domain_realization":"Vary equivalence definitions, subgroup thresholds, error costs, and downstream-performance assumptions."},{"component":"Pilot or Sunset Path","status":"direct","domain_realization":"Use retrospective shadow scoring followed by a small live pilot that expires absent affirmative renewal."}],"mechanism_dispositions":[{"slug":"congestion_or_capacity_pricing_adjustment","disposition":"incompatible","contribution_type":"NONE","adaptation_or_rejection":"The problem is classification, not peak-load rationing.","counterfactual_removal":"Removal leaves the causal chain unchanged."},{"slug":"cost_benefit_assessment_protocol","disposition":"selected_supporting","contribution_type":"TEST_DESIGN","adaptation_or_rejection":"Compare corrected classifications with validation, review, and misclassification costs under sensitivity analysis.","counterfactual_removal":"The rubric could still change, but its welfare and distributional case would be materially less reviewable."},{"slug":"distortion_reduction_review","disposition":"selected_load_bearing","contribution_type":"CORE_CAUSAL","adaptation_or_rejection":"Diagnose whether canonical-form exclusion, rather than genuine task failure, causes the blocked access.","counterfactual_removal":"Without this split, the intervention cannot distinguish avoidable exclusion from valid proficiency protection."},{"slug":"impact_assessment_table","disposition":"selected_supporting","contribution_type":"SAFETY_GUARDRAIL","adaptation_or_rejection":"Record subgroup gains, losses, protected constraints, and halt indicators.","counterfactual_removal":"Aggregate accuracy could conceal concentrated harm and weaken the safety gate."},{"slug":"matching_improvement_program","disposition":"considered_rejected","contribution_type":"NONE","adaptation_or_rejection":"Candidates and institutions already meet through the assessment; pairing is not the binding failure.","counterfactual_removal":"Removal leaves the proposed classification repair intact."},{"slug":"permit_or_approval_streamlining","disposition":"considered_rejected","contribution_type":"NONE","adaptation_or_rejection":"The target is a substantive scoring criterion, not duplicative approval processing.","counterfactual_removal":"Removal avoids mislabeling validity checks as procedural red tape."},{"slug":"price_control_redesign","disposition":"incompatible","contribution_type":"NONE","adaptation_or_rejection":"There is no ceiling, floor, or subsidy producing the classification error.","counterfactual_removal":"Removal has no causal effect."},{"slug":"quota_or_allocation_rule_review","disposition":"selected_load_bearing","contribution_type":"CORE_CAUSAL","adaptation_or_rejection":"Adapt the eligibility-formula review to a scoring threshold that allocates placement or certification.","counterfactual_removal":"Without changing the coarse allocation rule, valid variants remain excluded."},{"slug":"regulatory_simplification_pilot","disposition":"selected_load_bearing","contribution_type":"SAFETY_GUARDRAIL","adaptation_or_rejection":"Adapt its walled-off, monitored, expiring trial to assessment governance.","counterfactual_removal":"Absent a bounded reversible trial, uncertain validity and fairness risks hard-gate responsible implementation."},{"slug":"sunset_clause_review","disposition":"considered_rejected","contribution_type":"NONE","adaptation_or_rejection":"A separate sunset mechanism is redundant because the selected pilot already expires automatically.","counterfactual_removal":"The pilot's expiry and rollback remain intact."},{"slug":"tariff_fee_or_toll_redesign","disposition":"incompatible","contribution_type":"NONE","adaptation_or_rejection":"Assessment fees are not the hypothesized source of variant rejection.","counterfactual_removal":"Removal leaves the causal chain unchanged."}],"causal_chain":["A canonical-form-only rubric treats some task-adequate variants as errors.","Those errors lower scores or cross placement and certification thresholds.","Candidates incur denial, misplacement, retesting, or appeal despite adequate communicative performance.","A validated equivalence rubric accepts only variants shown to preserve the assessed function.","Correct classifications increase while reliability and genuine proficiency protections are monitored.","Preset expiry and rollback contain validity, gaming, and distributional harms."],"baseline":"Use the existing canonical-form rubric, ordinary rater calibration, and case-by-case appeals.","nearest_rival":"Improve rater training and appeals while leaving the set of accepted forms unchanged; this addresses inconsistent application but not exclusion encoded in the rubric.","authority_safety":{"affected_parties":["candidates","dialect communities","raters","assessment owners","score-using institutions","people relying on certified communication ability"],"decision_authority":"The assessment owner, subject to its psychometric, accreditation, legal, and affected-party review obligations.","authorized_first_step":"Shadow-score a stratified sample of archived responses using a preregistered equivalence rubric, then run one limited live cohort only if reliability and validity gates pass.","excluded_actions":["automatic acceptance of every noncanonical form","lowering task-relevant proficiency thresholds","retroactive score changes without due process","system-wide deployment before validation","using aggregate gains to waive subgroup safeguards"],"halt_rollback":"Stop live use and revert to the prior rubric if inter-rater reliability, downstream validity, communication-failure rates, or subgroup false classifications breach preregistered bounds; preserve no-fee appeal and correction for affected candidates."}},"negative_tests":{"strongest_counterevidence":"Archived and downstream data may show that disputed variants are rarely penalized, that penalties reflect task-relevant ambiguity rather than form alone, or that canonical-form performance predicts required communication outcomes materially better.","analogy_break":"Language assessment is not voluntary exchange: form can be part of the construct being certified, equivalence is context-sensitive, and expanding accepted variants may change rather than merely unblock the measured activity.","failure_condition":"The candidate fails if no separable class of communicatively adequate but rubric-rejected variants can be identified without weakening the assessment's stated construct.","problem_falsifier":"After controlling for task performance and rater effects, variant status does not independently predict scoring, placement, retesting, or appeal outcomes, or any observed difference is substantively negligible.","intervention_falsifier":"The expanded rubric does not reduce validated false rejections, or it materially worsens reliability, downstream communication performance, administrative burden, or protected-group error rates relative to baseline and rater-training rival.","risks":["mistaking genuine semantic or pragmatic differences for equivalence","reduced inter-rater reliability","strategic exploitation of ambiguous forms","tokenistic inclusion of only well-documented varieties","new subgroup disparities","construct drift where standardized-form mastery is itself required","privacy or stigmatization from dialect labeling"]},"null_rationale":null,"classification":{"candidate_kind":"TESTABLE_CONJECTURE","prior_art_status":"UNSEARCHED","evidence_maturity":"HYPOTHESIS"},"revision_change_log":{"revision_kind":"ORIGINAL","prior_problem_id":null,"prior_causal_lever_id":null,"problem_changed":false,"causal_lever_changed":false,"conceptual_changes":[],"operational_changes":[],"repairs_addressed":[]},"confidence":0.82,"generator_notes":"Closed-book structural inference from the supplied packet. Domain-specific empirical propositions are labeled HYPOTHESIS or INFERENCE; prior art was not searched."}