{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__human_computer_interaction:P4:v0","cell_id":"computability_boundary_mapping__human_computer_interaction","search_queries":["site:w3.org adaptive personalization semantics explanation accessibility specification personalized interface","adaptive user interfaces explanation transparency user study personalization explanations HCI","program slicing dynamic slicing execution trace feature influence explanation prior art","EU AI Act transparency deployers logs human oversight official regulation explanation","site:nist.gov AI RMF explainability interpretability transparency logs users official","local versus global explanations distinction primary paper explainable AI local global scope","program input influence undecidable semantic dependence variable influence Rice theorem paper","adaptive user interface transparency gap explanations study users why adaptation happened","Lim Dey Avrahami Why and Why Not explanations improve intelligibility context-aware intelligent systems ACM 2009","dynamic slicing specific execution trace original paper Korel Laski 1988 ACM","W3C PROV-O Recommendation provenance activities entities agents standard","model checking finite state systems exhaustive verification official handbook CMU","site:eur-lex.europa.eu Regulation EU 2024/1689 Article 13 transparency instructions deployers logs explanation decisions","site:eur-lex.europa.eu GDPR Article 5 data minimisation integrity confidentiality official","site:nist.gov NIST AI 100-1 explainable interpretable transparency accountable documentation human oversight PDF","Henry Gordon Rice classes recursively enumerable sets decision problems 1953 PDF","Rice theorem original 1953 Transactions AMS semantic properties programs pdf","halting problem undecidable official textbook open access computation theorem","site:arxiv.org dynamic program slicing execution trace paper pdf","site:edu \"Rice's theorem\" \"nontrivial semantic\" program undecidable","site:stanford.edu encyclopedia Rice theorem computability semantic properties programs"],"sources":[{"source_id":"S1","title":"WAI-Adapt Explainer","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/adapt/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023-01-03","accessed_at":"2026-08-03","claims_supported":["Personalized interfaces are an established accessibility use case involving changes to controls, symbols, help, complexity, and layout.","W3C identifies users with cognitive and learning disabilities, authors, user-agent developers, and accessibility implementers as stakeholders in adaptable interfaces.","The document is a Group Draft Note rather than an endorsed W3C Recommendation."]},{"source_id":"S2","title":"Why and Why Not Explanations Improve the Intelligibility of Context-Aware Intelligent Systems","publisher":"Association for Computing Machinery / Carnegie Mellon University","url":"https://www.cs.cmu.edu/~byl/publications/lim_chi09.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2009-04-04","accessed_at":"2026-08-03","claims_supported":["Context-aware systems using implicit inputs and complex rules can be difficult for users to understand, contributing to mistrust, misuse, or abandonment.","A controlled study with 211 participants distinguished Why, Why Not, What If, and How To explanation questions.","Why and Why Not explanations improved task performance and understanding relative to no explanation, while effects differed by explanation type and imposed time or cognitive costs."]},{"source_id":"S3","title":"Four Principles of Explainable Artificial Intelligence","publisher":"National Institute of Standards and Technology","url":"https://www.govinfo.gov/content/pkg/GOVPUB-C13-7848d8b02b0f9467e09670d6f7531430/pdf/GOVPUB-C13-7848d8b02b0f9467e09670d6f7531430.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2021-09","accessed_at":"2026-08-03","claims_supported":["NIST identifies explanation, meaningfulness, explanation accuracy, and knowledge limits as separate principles.","An explanation can be understandable yet fail to reflect the system's actual process.","Systems should identify out-of-domain or unreliable cases rather than issue potentially misleading answers.","NIST separately describes local, counterfactual, and global explanation scopes and states that explanation needs vary by audience and purpose."]},{"source_id":"S4","title":"Local vs. Global Interpretability: A Computational Complexity Perspective","publisher":"Proceedings of Machine Learning Research","url":"https://proceedings.mlr.press/v235/bassan24a.html","source_class":"PRIMARY_RESEARCH","publication_date":"2024-07-21","accessed_at":"2026-08-03","claims_supported":["Local and global interpretability are established as mathematically distinct explanation perspectives.","The computational difficulty of explanation depends on explanation scope and model class.","The paper analyzes selected machine-learning model classes and complexity, not arbitrary executable adaptation programs or the candidate's undecidability claim."]},{"source_id":"S5","title":"Dynamic Slicing by On-demand Re-execution","publisher":"arXiv","url":"https://arxiv.org/abs/2211.04683","source_class":"PRIMARY_RESEARCH","publication_date":"2022-11-09","accessed_at":"2026-08-03","claims_supported":["Dynamic slicing identifies statements affecting a value on a specific execution path using observed data and control dependencies.","Trace-based slicing is mature prior art for actual-execution provenance but is inherently tied to particular executions.","The evaluated implementation had scalability, library-modeling, determinism, and possible recall limitations, illustrating implementation risks for production provenance."]},{"source_id":"S6","title":"PROV-O: The PROV Ontology","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/prov-o/","source_class":"STANDARD","publication_date":"2013-04-30","accessed_at":"2026-08-03","claims_supported":["PROV-O is an endorsed W3C Recommendation for representing provenance.","It provides Entity, Activity, Agent, usage, generation, derivation, responsibility, and qualified-influence concepts suitable for an expandable evidence record.","PROV-O represents provenance relationships but does not itself prove causal necessity, global influence, or non-influence."]},{"source_id":"S7","title":"Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence","publisher":"European Union, EUR-Lex","url":"https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32024R1689","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-07-12","accessed_at":"2026-08-03","claims_supported":["For high-risk AI systems, providers must supply concise, complete, correct, clear, accessible, and comprehensible information about capabilities, limitations, interpretation, oversight, and relevant logging mechanisms.","The regulation identifies providers and deployers as responsible actors and requires effective human oversight for high-risk systems.","Deployers must retain automatically generated logs under specified conditions, subject to applicable personal-data law.","These obligations create conditional authorizer pull but do not establish that every adaptive interface is a regulated high-risk AI system."]},{"source_id":"S8","title":"Recursive Functions","publisher":"Stanford Encyclopedia of Philosophy","url":"https://plato.stanford.edu/archives/spr2026/entries/recursive-functions/","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2026-Spring","accessed_at":"2026-08-03","claims_supported":["Rice's theorem states that every non-trivial semantic index set is undecidable.","The theorem's proof transfers undecidability through an effective reduction.","Computably enumerable sets support one-sided enumeration, while decidability requires an always-terminating correct algorithm.","Applying the theorem to the candidate still requires a checked formalization of the adaptation language, feature-isolation relation, output semantics, and reduction."]}],"problem_evidence":{"support":"STRONG","rationale":"Adaptive and context-aware interfaces visibly exist, including accessibility-oriented personalization (S1). A primary HCI study documents user difficulty reasoning about implicit-input, rule-driven behavior and reports measurable understanding and task-performance benefits from appropriately chosen explanations (S2). NIST independently identifies explanation accuracy, audience meaningfulness, and declaration of knowledge limits as distinct requirements, directly supporting the risk of persuasive explanations that exceed their evidence (S3). However, no source measures how often adaptive-interface products currently collapse trace, bounded, and global claims, so prevalence and realized harm remain unmeasured.","source_ids":["S1","S2","S3"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable stakeholders include adaptive-interface users and accessibility implementers described by W3C (S1), product developers and end users studied in context-aware systems (S2), and—conditionally—EU high-risk-AI providers, deployers, and oversight personnel who have binding transparency, limitation, interpretation, and logging duties (S7). This establishes expressed institutional need and an authorizer class, but not a named product team, budget owner, procurement commitment, or evidence that the candidate component falls within high-risk-AI scope.","source_ids":["S1","S2","S7"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"NIST explanation-accuracy and knowledge-limits principles","similarity":"Already requires explanations to reflect the actual system process, suit the intended audience, and expose conditions where answers are unreliable or out of scope; it also distinguishes local, counterfactual, and global explanations.","remaining_difference":"It is principle-level guidance, not an executable adaptive-rule query router with formal finite-fragment certificates, replayable influence witnesses, and separate UNKNOWN_GLOBAL_INFLUENCE and OUT_OF_MODEL_DEPENDENCY states.","source_ids":["S3"]},{"name":"Lim-Dey-Avrahami context-aware intelligibility framework","similarity":"Separates Why, Why Not, What If, and How To questions for adaptive/context-aware behavior and empirically evaluates their effects on users.","remaining_difference":"Its studied rule systems are bounded experimental models; it does not classify the computability of unrestricted program-level global influence or prohibit exact negative answers outside an enforceable fragment.","source_ids":["S2"]},{"name":"Dynamic program slicing and W3C PROV-O","similarity":"Together provide established methods and a standard vocabulary for recording trace-specific dependencies, entities, activities, derivations, and responsibility.","remaining_difference":"They establish actual-execution provenance, not global counterfactual influence or non-influence; the candidate's distinctive contribution is to prevent provenance from being promoted to a stronger quantifier.","source_ids":["S5","S6"]},{"name":"Formal local/global interpretability analysis plus Rice's theorem","similarity":"Prior work already distinguishes explanation scope mathematically, examines its computational cost, and establishes undecidability for non-trivial semantic properties of unrestricted programs.","remaining_difference":"No reviewed source applies that boundary to adaptive-interface feature influence and packages restricted exact analysis, unrestricted witness search, trace provenance, external-dependency disclosure, and user-facing status labels as one governed contract.","source_ids":["S4","S8"]}],"distinctive_claim_remaining":"For executable adaptive-interface rules, a versioned router can preserve explanation quantifiers operationally: trace evidence is labeled only as trace provenance; exact global non-influence is issued only for mechanically enforced finite-total rules; unrestricted rules yield a replayable positive witness or explicit unknown; and undeclared external dependencies remain a distinct status. The claim is falsified by any mislabeled seeded case, bypassable fragment check, unreplayable witness, starvation of a finite witness, hidden external dependency, or user interpretation of an incomplete state as exact non-influence.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Core primitives are credible: dynamic slicing can extract execution-specific dependencies (S5); PROV-O can structure provenance and responsibility records (S6); finite-model enumeration is constructive once domains, state, and termination are truly finite; and Rice's theorem supports refusing a total semantic analyzer for unrestricted programs (S8). NIST supplies a compatible knowledge-limit and explanation-accuracy contract (S3). The integrated router, enforceable grammar, reduction, exact decider, fairness of dovetailing, feature-isolation semantics, external-dependency detection, privacy controls, latency, and user comprehension have not been implemented or independently tested. Dynamic-slicing evidence also reports modeling, nondeterminism, recall, and scaling limitations.","source_ids":["S3","S5","S6","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"The intervention could prevent misleading explanations, bad governance conclusions, and inaccessible adaptive behavior, and explanation quality has measured effects on understanding and task performance. Candidate-specific prevalence, severity, and realized impact are absent.","source_ids":["S1","S2","S3"]},"stakeholder_pull":{"score":3,"rationale":"W3C documents accessibility demand and EU law creates conditional transparency and logging obligations for identifiable providers and deployers. No named adopter, committed owner, or budget has been verified.","source_ids":["S1","S7"]},"incremental_advantage":{"score":3,"rationale":"The router offers a useful guarantee upgrade over reason strings, logs, or sampled counterfactuals by keeping claim scopes distinct. Most constituent ideas are established, and comparative operational benefit is untested.","source_ids":["S2","S3","S5","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"No exact integrated analogue was found, but the proposal closely composes established local/global explanation taxonomies, knowledge-limit guidance, dynamic slicing, provenance standards, and computability theory. World novelty remains unmeasured.","source_ids":["S2","S3","S4","S5","S6","S8"]},"technical_implementability":{"score":4,"rationale":"An offline fixture implementation is technically credible using tracing, a finite DSL, enumeration, evidence records, and explicit abstention. Production soundness is threatened by nondeterminism, library or service modeling, aliases, grammar escape hatches, state explosion, and trace sensitivity.","source_ids":["S5","S6","S8"]},"adoption_authority_feasibility":{"score":4,"rationale":"A product owner can authorize an offline shadow audit, while design, accessibility/data, and formal-review owners can gate semantics and evidence. EU obligations provide a credible external authorization context for some systems. Production legal scope remains conditional.","source_ids":["S1","S7"]},"evidence_readiness":{"score":3,"rationale":"The candidate specifies concrete fixtures, statuses, halt conditions, and falsifiers, and established sources provide comparators. No implementation, checked proof, fixture results, human study, or proprietary workflow evidence exists.","source_ids":["S2","S3","S5","S8"]},"safety_net_benefit":{"score":4,"rationale":"Explicit unknown, out-of-scope, tool-failure, and external-dependency states can prevent false exact negatives and support oversight. Benefits depend on labels being technically correct and understood; sensitive provenance may create privacy and security risk.","source_ids":["S3","S6","S7"]},"scalability":{"score":3,"rationale":"Trace provenance and routing can be reused across components, but exact enumeration can suffer state explosion, unrestricted searches can overproduce unknown, and dynamic slicing reports execution-size, library-modeling, nondeterminism, and recall constraints.","source_ids":["S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One nonproduction component; frozen semantics; approximately 30-50 synthetic and recorded fixtures; prototype router; independent paper proof review; security review of audit artifacts; and a small preregistered label-comprehension study.","confidence":"LOW","assumptions":["Approximately 3-8 person-weeks across formal methods, personalization engineering, interaction research, and security/privacy review.","Existing nonproduction execution harness and representative rule examples are available.","No production integration, procurement, or regulated conformity assessment is included.","This is a bottom-up resource-equivalent estimate; none of the eight sources provides a project price or labor-rate benchmark."],"source_ids":["S2","S3","S5","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build a production-quality finite-rule checker and exact analyzer for one component, instrument trace provenance, implement signed/versioned evidence records, integrate external-dependency declarations, and complete independent formal and privacy/security reviews.","confidence":"LOW","assumptions":["One adaptation engine and one outcome predicate are in scope.","Existing identity, logging, UI-component, and CI infrastructure can be reused.","The rule language can be restricted without replacing the personalization platform.","No organization-wide migration or high-risk-AI conformity assessment is included."],"source_ids":["S5","S6","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production rollout for one product surface with hardened observability, access control, retention policy, support and contestability workflow, accessibility and interaction validation, load testing, incident rollback, training, and staged deployment.","confidence":"LOW","assumptions":["A cross-functional team works for roughly one to three quarters.","Sensitive traces require controlled storage and review interfaces.","The launch includes integration and governance work but not rebuilding arbitrary upstream models or services.","Costs rise materially if regulated certification, multilingual UX, or multiple adaptation engines are required."],"source_ids":["S1","S3","S5","S6","S7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Ongoing rule-fragment and dependency review, proof and fixture regression, monitoring of unknown rates and witness replay, security/privacy audits, user-support escalation, documentation, and reclassification after rule or model changes.","confidence":"LOW","assumptions":["Approximately 0.5-2.0 full-time-equivalent annual effort plus compute and audit overhead.","One product and a modest rate of rule-language changes are covered.","Large-scale exhaustive searches are budget-capped and return unknown rather than consuming unbounded resources.","No external source directly verifies this recurring-cost range."],"source_ids":["S3","S5","S6","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent official and primary sources establish real adaptive-interface use, user intelligibility problems, measurable explanation effects, and the need for accurate, scope-aware explanations, although prevalence of the exact failure mode is unknown.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"EU regulators are an identifiable authorizer for high-risk-AI transparency and logging, and providers/deployers are identifiable adopting actors; W3C identifies accessibility implementers. Applicability to a particular product must still be established.","source_ids":["S1","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The scope-preservation claim can be tested with seeded trace, finite, unrestricted, external-dependency, timeout, and failure fixtures and with user interpretation measures against reason-code and unlabeled-explanation comparators.","source_ids":["S2","S3","S5","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A one-component offline audit with frozen semantics, finite fixtures, independent proof review, explicit comparators, numerical continuation thresholds, and enumerated falsifiers is bounded in scope and cost.","source_ids":["S2","S3","S5","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The proposed first step is nonproduction, cannot alter user outcomes, can use synthetic or minimized authorized traces, and has explicit halt and rollback conditions. Privacy, trace access, and semantic approval remain required controls but are not unavoidable stops for the offline step.","source_ids":["S3","S6","S7"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The four scopes and assumptions are bounded and internally plausible, but no direct quote, labor-rate benchmark, infrastructure inventory, or adopter estimate was found among the eight sources. The bands are resource-equivalent planning estimates, not externally validated budgets.","source_ids":["S5","S6","S7"]}},"next_evidence_step":"Run a preregistered offline shadow study on exactly one nonproduction adaptive component. Freeze the rule semantics, feature-isolation relation, context domains, output predicate, status alphabet, finite-fragment grammar, external-dependency declarations, and resource bounds. Build 30-50 seeded fixtures spanning known trace participation, finite global relevance and irrelevance, delayed influence, a nonterminating candidate scheduled before a finite witness, bounded non-finding, feature aliasing, nondeterminism, undeclared services, missing logs, and tool failure. Independently check the reduction and finite decider. Compare (A) current reason codes/log viewer, (B) the same evidence presented without scope labels, and (C) the proposed router and expandable evidence record with at least 24 representative users, designers, support staff, and reviewers. Continue only if every exact finite verdict is correct, every witness replays, no nonterminating job starves a later witness, zero incomplete cases are rendered as non-influence, all external dependencies are disclosed, no unauthorized trace data are exposed, at least 80% of participants correctly identify each label's quantifier, and condition C improves scope comprehension by at least 20 percentage points over both comparators without a material task-time or accessibility regression. Falsify the intervention if fragment membership is bypassable, any seeded exact negative is false, a valid witness is rejected or unreplayable, an external dependency is hidden, unknown and failure states collapse, or the labeled interface fails the comprehension threshold.","blocking_evidence":["No measured prevalence or severity of trace-to-global overclaiming in deployed adaptive interfaces.","No named adopter, product owner commitment, budget, or confirmed regulatory classification for a target component.","No independently checked reduction for the exact feature-isolation and output semantics.","No implemented finite-fragment checker or evidence that its membership boundary cannot be bypassed.","No fixture results for classification accuracy, termination, fair scheduling, witness replay, external dependencies, or tool failures.","No field evidence that intended users understand the status labels and quantifiers or find them useful.","No target-specific privacy, retention, access-control, nondeterminism, feature-alias, or external-service assessment.","No externally validated cost quote or resource inventory.","World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search found established adaptive-interface explanation research, local/global explanation distinctions, explanation-accuracy and knowledge-limit guidance, trace-specific dynamic slicing, provenance standards, and general computability limits. It did not establish whether the particular adaptive-interface contract—finite-fragment exact influence, unrestricted replayable witnesses or explicit unknown, trace-only provenance, external-dependency status, and versioned user-facing routing—has been published, patented, commercialized, or deployed. World novelty, patentability, freedom to operate, market size, and realized impact are explicitly unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Implement the one-component frozen fixture harness and query router.","Obtain independent review of the exact reduction, finite decider, and fragment-enforcement boundary.","Demonstrate perfect classification on seeded exact cases, replayable witnesses, fair scheduling, and distinct incomplete states.","Measure label-scope comprehension against reason-code and unlabeled-evidence comparators with representative stakeholders.","Complete target-specific privacy, trace-access, retention, nondeterminism, feature-alias, and external-dependency reviews.","Obtain a named adopter decision owner, applicability determination, infrastructure inventory, and bottom-up cost estimate."],"reason":"Bounded web research establishes that the problem is credible, the component is technically plausible, and its ingredients are mostly adjacent established practice. It cannot establish proof correctness for the candidate's exact semantics, fragment enforceability, production trace fidelity, witness replay, user comprehension, workflow fit, privacy safety, adopter commitment, or actual cost. Those questions require implementation, proprietary workflow evidence, and live human testing, so further web search cannot clear the remaining gates."},"proposal_index":4}