{"schema_version":1,"assessment_id":"eoa_inverse_innovation_exp03_opportunity320_20260801","source_experiment_id":"eoa_inverse_innovation_exp03_full320_20260801","cell_id":"computability_boundary_mapping__education_pedagogy","archetype_slug":"computability_boundary_mapping","domain_slug":"education_pedagogy","title":"Explicit Computability Boundaries and UNKNOWN Routing for Program Autograding","opportunity_summary":"Audit whether a programming autograder promises universal terminating pass/fail judgments, then restrict exact grading to an enforced decidable fragment and route out-of-scope, timeout, analyzer-failure, and unresolved cases to labeled UNKNOWN or human review. The candidate could prevent unsupported fail verdicts, but the sealed packet does not establish that any concrete course has the alleged problem.","adopter_authorizer":"A course lead may authorize a non-grading shadow study; live grading-policy changes require the institution's normal faculty assessment and appeals authority, with accessibility and appeals staff involved.","scores":{"meaningful_impact":{"score":4,"rationale":"If the stated collapse of timeout or unsupported cases into failure occurs, the intervention addresses consequential grades, progression, appeals, explainability, and equity. The number of affected students and frequency of such errors are unsupported."},"stakeholder_pull":{"score":2,"rationale":"The candidate identifies affected roles and plausible burdens but provides no named course, institution, procurement process, expressed demand, or evidence that stakeholders recognize this as a priority."},"incremental_advantage":{"score":4,"rationale":"Unlike the stated rival of improving tests, capacity, and timeouts while retaining Boolean universality, the proposal changes the grading contract itself by enforcing an exact region and preserving UNKNOWN and system-failure states. Its advantage disappears if the current grader already does this."},"distinctiveness_plausibility":{"score":3,"rationale":"The composition of formal scope classification, checked boundary evidence, enforced fragments, labeled routing, and reclassification triggers is specific and testable, but prior art is explicitly unsearched and no novelty claim is supportable."},"technical_implementability":{"score":3,"rationale":"A single assignment with a formal executable property and bounded fragment appears technically tractable, while semantic rubric ambiguity, fragment-membership errors, proof mistakes, sandbox integration, and boundary-gaming create substantial implementation difficulty."},"adoption_authority_feasibility":{"score":4,"rationale":"The packet clearly assigns shadow-pilot authority to a course lead and reserves policy changes for existing faculty and appeals processes. Feasibility remains conditional on finding a participating course and securing approved access to student code and logs."},"evidence_readiness":{"score":4,"rationale":"The candidate supplies independent problem and intervention falsifiers, a bounded shadow comparison, disagreement review, subgroup monitoring, halt criteria, and rollback. It lacks a concrete target, baseline observations, and validated rubric formalization."},"safety_net_benefit":{"score":5,"rationale":"The central mechanism preserves uncertainty rather than converting it to failure, retains human review and the prior appeals path, forbids grade changes during testing, and specifies immediate halt conditions for refuted exactness or UNKNOWN-to-FAIL leakage."},"scalability":{"score":2,"rationale":"Each assignment may require separate property formalization, fragment enforcement, proof review, validity checks, and human escalation capacity. Diverse languages and pedagogical objectives limit reuse, and no cross-course operating evidence is provided."}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Audit one identified course, one recent programming assignment, its published grading contract, accepted language, interface states, relevant logs, and appeals handling against the problem falsifier.","confidence":"MODERATE","assumptions":["Approved internal access to specifications and de-identified or appropriately governed logs is available.","The audit requires instructor, autograder developer, assessment, and limited compliance effort.","No software deployment or grade changes occur."]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Create and independently review a shadow-mode implementation for one assignment, including formal property definition, fragment checker, exact/UNKNOWN/failure routing, logging, and disagreement review.","confidence":"LOW","assumptions":["The targeted property is formalizable and primarily concerns program behavior.","Existing autograder and sandbox components can be extended rather than replaced.","Student code remains within approved systems.","The band excludes broad multi-course rollout."]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Move a validated one-assignment design into authorized grading operations for one course, including governance approval, gradebook integration, staff training, accessibility and appeals procedures, monitoring, and launch evaluation.","confidence":"LOW","assumptions":["Shadow evidence supports correctness and manageable UNKNOWN volume.","Normal faculty assessment authority approves the policy change.","Human-review capacity is available without building a new central service.","No major procurement or platform replacement is required."]},"annual_recurring":{"band_2026_usd":"10K_TO_50K","scope":"Operate and maintain the mechanism for one course, including review of UNKNOWN cases, appeals support, monitoring, and revalidation after material rubric, language, or platform changes.","confidence":"LOW","assumptions":["Use remains limited to a small number of assignments in one course.","UNKNOWN volume is low enough for existing staff plus bounded additional effort.","Material curriculum changes trigger revalidation rather than automatic reuse.","Program-wide scaling could exceed this band."]}},"research_burden":"HIGH","earliest_credible_horizon":"0_TO_3_MONTHS","pipeline_gates":{"recognizable_externally_supportable_problem":{"status":"UNCERTAIN","reason":"The problem has observable audit criteria and meaningful consequences, but the packet supplies no evidence that a concrete course makes the unrestricted claim or collapses exceptional states into failure."},"identifiable_adopter_or_authorizer":{"status":"YES","reason":"The course lead is identified as shadow-study authorizer, while faculty assessment and appeals authorities control grading-policy changes."},"distinct_testable_incremental_claim":{"status":"YES","reason":"The proposal claims that enforced exact-region boundaries plus labeled UNKNOWN and human routing will reduce unsupported fail classifications relative to a test-suite-and-timeout Boolean grader; the shadow comparison can falsify this claim."},"bounded_next_evidence_step":{"status":"YES","reason":"A one-course, one-assignment specification-and-log audit is bounded, non-grading, and can falsify the alleged problem before any implementation."},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first audit and proposed shadow test do not alter grades, keep code in approved systems, preserve appeals, require independent review, and include explicit halt and rollback conditions."},"implementation_cost_scope_and_range":{"status":"UNCERTAIN","reason":"A one-assignment scope supports broad resource bands, but actual cost depends on the unidentified platform, rubric formalizability, data governance, submission volume, integration burden, and human-review rate."}},"blocking_evidence":["Whether a concrete target grader actually quantifies over an open-ended language or property and records timeout, unknown, out-of-scope, or system failure as incorrect.","Whether the targeted rubric can be formalized without narrowing or changing the intended learning objective.","Whether fragment membership and every purportedly exact verdict remain correct under independent review.","Whether labeled routing reduces unsupported failures while keeping UNKNOWN volume, feedback delay, workload, and subgroup disparities acceptable.","Whether approved access, governance authority, and sustained human-review capacity exist at a participating course."],"next_evidence_step":"Within one identified course, conduct a read-only audit of one recent programming assignment and one term of relevant records: compare the published contract, mechanically accepted inputs, interface state taxonomy, timeout and analyzer logs, and appeals flow with the problem falsifier. Conclude no-go for this mechanism if all accepted inputs are already in an enforced decidable class and incorrect, timeout, UNKNOWN, out-of-scope, and system failure remain distinct through grading and appeals; do not change grades or deploy new routing.","research_questions":["Does the target course make an unrestricted semantic guarantee, explicitly or operationally, and how are exceptional states represented through the gradebook and appeals process?","Can the assignment's graded property and accepted program class be formalized without invalidating the intended assessment construct?","What independently checkable evidence supports correctness within the proposed exact region and correct fragment-membership classification?","Compared with current verdicts, how many cases would become exact, UNKNOWN, out-of-scope, or system failure, and are disagreement and subgroup patterns acceptable?","What human-review workload, turnaround time, governance, accessibility, and appeals capacity would labeled routing require?","Do existing autograding or program-analysis practices already implement the proposed restriction and explicit-UNKNOWN composition? "],"recommendation":"VALIDATE_PROBLEM_FIRST","uncertainty_constraints":["Closed-book assessment: no external evidence was used.","Problem prevalence, stakeholder demand, realized educational impact, and market size are unmeasured.","No institution, course, assignment, autograder architecture, submission volume, or baseline error rate is identified.","Prior art is explicitly unsearched, so distinctiveness and novelty cannot be established.","Cost bands are resource-equivalent scenario ranges based on a one-course scope, not estimates for a known implementation.","Program-property correctness does not establish authorship, understanding, intent, creativity, or broader pedagogical mastery."],"closed_book_prior_art_boundary":"The sealed candidate labels prior art as UNSEARCHED and makes no novelty or prevalence claim. This assessment therefore treats formally restricted autograders, explicit UNKNOWN verdicts, and human-review routing as potentially existing practices and does not infer distinctiveness from their absence in the packet."}