{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__religious_studies_theology:P1:v0","cell_id":"predictive_residual_processing__religious_studies_theology","search_queries":["digital humanities ritual liturgy computational analysis repeated texts variants corpus","digital liturgy database compare ritual texts variants project","religious studies corpus annotation ritual text digital humanities","TEI critical apparatus CollateX variant texts official documentation","CollateX official documentation collation variants witnesses","CARE Principles Indigenous Data Governance official GIDA collective benefit authority control ethics","CATMA official digital humanities annotation platform provenance collaborative","ritual studies digital corpus repeated ritual manuscripts variants computational comparison","site:aclanthology.org/2026.nlp4dh Vedic ritual computational analysis ritual content versions","\"Vedic ritual\" \"computational analysis\" 2026 NLP4DH","\"Repetition analysis function\" Avestan ritual manuscript DOI","site:ititerr.it T-ReS CRITERION GNORM critical editions religious studies"],"sources":[{"source_id":"S1","title":"Arthur Westwell: Digital Techniques for Presenting Liturgical Texts and Building a Database of Carolingian Pontificals","publisher":"Trier Center for Digital Humanities, University of Trier","url":"https://tcdh.uni-trier.de/de/arthur-westwell-digital-techniques-presenting-liturgical-texts-and-building-database-carolingian","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Liturgical manuscripts are nonuniform and contain interventions that explain, edit, or reinterpret ritual texts.","Printed editions can obscure variance and create a false appearance of uniformity.","A named researcher and university center sought a digital edition facilitating comparison of approximately 40 manuscripts while exposing every manuscript variant."]},{"source_id":"S2","title":"Corpus of Hittite Festive Rituals: Description","publisher":"Academy of Sciences and Literature Mainz","url":"https://www.adwmainz.de/en/research/projects/corpus-der-hethitischen-festrituale/description.html","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["The ritual corpus contains more than 10,000 fragments and presents volume, fragmentation, reconstruction, and indexing challenges requiring long-term cooperative research and tailored IT infrastructure.","The project uses an extensively annotated digital corpus, searchable metadata, synoptic manuscript presentation, master text, translations, annotations, and source photographs.","The Mainz Academy and partner universities are identifiable potential institutional adopters or authorizers for bounded research tooling."]},{"source_id":"S3","title":"WP3 – T-ReS – Toolkit for Religious Studies","publisher":"Italian Strengthening of the ESFRI RI RESILIENCE","url":"https://www.itserr.it/itserr-its-work-packages/wp3-t-res/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["T-ReS was explicitly designed to address concrete scholarly needs in religious studies: critical editions and analysis of complex normative corpora.","CRITERION already manages texts, variants, witnesses, multilevel annotations, metadata, collaboration, TEI-XML export, and publication workflows.","GNORM already applies preprocessing, entity recognition, exact and approximate string matching, reference normalization, and diachronic visualization to stratified religious corpora."]},{"source_id":"S4","title":"TEI Guidelines, Chapter 13: Critical Apparatus","publisher":"Text Encoding Initiative Consortium","url":"https://guidelines.tei-c.de/en/html/TC.html","source_class":"STANDARD","publication_date":"2026-02-18","accessed_at":"2026-08-03","claims_supported":["TEI provides an established structured representation for variant readings, witnesses, apparatus entries, and interactive witness display.","The standard supports provenance-linked implementation but does not prescribe one universally correct text-critical method.","Identification and grouping of textual variation remain editorial judgments rather than purely mechanical operations."]},{"source_id":"S5","title":"CollateX Documentation","publisher":"CollateX / Interedition","url":"https://collatex.net/doc/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["CollateX aligns token sequences from multiple witnesses and represents commonality, differences, sequence, insertions, and omissions in tabular or variant-graph forms.","It accepts JSON or XML-related workflows and emits JSON, TEI, graph, and tabular outputs suitable for a prototype.","Transpositions and preservation of XML markup context have documented limitations, showing that structured ritual-event comparison is not turnkey."]},{"source_id":"S6","title":"Recurrence Analysis Function, a Dynamic Heatmap for the Visualization of Verse Text and Beyond","publisher":"Heidelberg University Publishing","url":"https://heiup.uni-heidelberg.de/catalog/view/345/501/128343/4400","source_class":"PRIMARY_RESEARCH","publication_date":"2018-04-12","accessed_at":"2026-08-03","claims_supported":["Computational recurrence analysis has already been applied to repetitive Avestan liturgical material.","Ritual manuscripts can encode intended performance through abbreviated repetitions, demonstrating that literal text alone may not preserve performative context.","Repetition visualization can reveal formulaic and text-genetic structures, making it close functional prior art for attention to recurring ritual units."]},{"source_id":"S7","title":"Quantifying Text Reuse Across Three Kṛṣṇa Yajurveda Recensions: Using Multi-Algorithm Computational Collation","publisher":"Association for Computational Linguistics","url":"https://aclanthology.org/2026.nlp4dh-1.5/","source_class":"PRIMARY_RESEARCH","publication_date":"2026-07-04","accessed_at":"2026-08-03","claims_supported":["A 2026 primary study computationally collated three recensions sharing substantial Vedic ritual content.","Five similarity algorithms found section-specific reuse patterns, including up to 93.5 percent overlap for one pair in one ritual section.","The results demonstrate both technical feasibility and the danger of assuming one uniform reuse model across ritual categories."]},{"source_id":"S8","title":"The CARE Principles for Indigenous Data Governance","publisher":"Global Indigenous Data Alliance","url":"https://www.gida-global.org/careprinciples","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Open-data practices can ignore power differentials, historical context, and Indigenous rights and interests.","Indigenous peoples assert control over uses of Indigenous data and knowledge.","People- and purpose-oriented governance, collective benefit, authority, responsibility, and ethics support the proposal's community-authority and restricted-material safeguards."]}],"problem_evidence":{"support":"MODERATE","rationale":"The underlying objects and stakes are visible: liturgical witnesses contain consequential nonuniformity that conventional presentation can obscure; live projects face corpora ranging from roughly 40 manuscripts to more than 10,000 fragments; and Vedic and Avestan research confirms extensive, context-dependent repetition. However, no opened source measures the proposal's specific bottleneck—researcher time lost rereading expected units—or shows that material variants are currently found late or inconsistently. The problem is therefore credible but its prevalence and magnitude in the proposed 60-record setting remain unmeasured.","source_ids":["S1","S2","S6","S7"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Identifiable institutions and scholars are already building relevant infrastructure. T-ReS explicitly responds to concrete religious-studies needs, while the Mainz Hittite project and Trier liturgical project require comparison, reconstruction, metadata, and variant access. These are credible adopters, authorizers, or partners, but none expresses demand for predictive collapse, residual prioritization, or willingness to pilot this particular interface.","source_ids":["S1","S2","S3"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"TEI critical apparatus plus CollateX","similarity":"Together they provide standardized witness representation and automated alignment of additions, omissions, substitutions, sequence, and variant graphs—the core descriptive comparison layer.","remaining_difference":"They do not establish a pre-observation probabilistic expectation, uncertainty-and-consequence queue, attention-savings claim, independent full-record audit, validated model updates, or automatic decompression on drift.","source_ids":["S4","S5"]},{"name":"ICoMa multi-algorithm collation of Vedic ritual recensions","similarity":"It computationally identifies and quantifies shared and divergent ritual text across recensions, using multiple algorithms and section-level comparison.","remaining_difference":"It measures text reuse for scholarship rather than collapsing predicted units in a reviewer workflow, reconstructing from a synchronized model plus residuals, or auditing suppressed context.","source_ids":["S7"]},{"name":"T-ReS CRITERION and GNORM","similarity":"These scholar-centered tools already manage witnesses, variants, annotations, metadata, collaboration, matching, NLP, stratified corpora, and religious-studies workflows.","remaining_difference":"The opened description does not show predictive residual encoding, randomized raw-record audits, collection-specific error budgets, or fallback triggered by reconstruction failure.","source_ids":["S3"]},{"name":"Recurrence Analysis Function for Avestan ritual material","similarity":"It uses computation to foreground repetitive and structurally informative passages in ritual manuscripts and can expose text-genetic patterns.","remaining_difference":"It is an exploratory visualization, not a governed predictor-residual-review loop with calibrated suppression and full-record fallback.","source_ids":["S6"]},{"name":"Hittite Festive Rituals and Carolingian Pontificals digital editions","similarity":"Both provide or seek synoptic comparison, extensive metadata, complete variant access, reconstruction, and links to source evidence in ritual corpora.","remaining_difference":"They retain full or synoptic presentation rather than allocating first-pass attention through a learned, versioned predictive baseline.","source_ids":["S1","S2"]}],"distinctive_claim_remaining":"For one authorized and narrowly bounded ritual-record collection, a versioned pre-observation sequence model plus structured residual interface, consequence-aware routing, synchronized reconstruction, independent full-record audits, and automatic full-display fallback will reduce total review time while remaining noninferior to full-record and nonpredictive-collation comparators for material-variant recall, reconstruction fidelity, context coverage, and protected-class handling. No opened source establishes that joint result.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"TEI, CollateX, CRITERION, GNORM, ICoMa, and existing ritual corpora show that witness encoding, alignment, metadata, sequence comparison, and scholar-facing tooling are technically feasible. A read-only retrospective prototype avoids altering sources and can retain direct source access. Important gaps remain: no source validates the proposed event-code ontology, semantic reconstruction tolerance, uncertainty calibration, model-update rule, subgroup audit design, or interface effect on interpretation. CollateX also documents transposition and markup limitations. Authority is collection-specific; CARE principles make clear that institutional possession or technical access does not substitute for community authority over culturally governed data.","source_ids":["S2","S3","S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Large, difficult ritual corpora and consequential local variation make improved review potentially valuable, but neither attention savings nor downstream scholarly benefit has been measured.","source_ids":["S1","S2","S7"]},"stakeholder_pull":{"score":3,"rationale":"Named religious-studies infrastructure projects express needs for critical editions, corpus analysis, comparison, and reconstruction, but no stakeholder requests this residual workflow or commits resources to it.","source_ids":["S1","S2","S3"]},"incremental_advantage":{"score":2,"rationale":"TEI, CollateX, T-ReS, ICoMa, recurrence analysis, and synoptic digital editions already cover much of the functional surface. The proposed advantage depends on untested attention savings and safety controls rather than a demonstrated capability gap.","source_ids":["S2","S3","S4","S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"The joint predictive-collapse, audit, and fallback architecture was not found in the eight sources, but its components are close extensions of established collation, apparatus, corpus, and comparison practices. World novelty remains unmeasured.","source_ids":["S3","S4","S5","S7"]},"technical_implementability":{"score":4,"rationale":"Standards, open tooling, structured outputs, existing annotated corpora, and recent multi-algorithm ritual collation support a retrospective prototype. Semantic event coding, transpositions, performance context, and calibration remain nontrivial.","source_ids":["S2","S3","S4","S5","S6","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"Universities, academy projects, corpus committees, and community stewards provide identifiable authorization pathways. Actual permission, data-use terms, and community authority must be confirmed corpus by corpus.","source_ids":["S1","S2","S3","S8"]},"evidence_readiness":{"score":3,"rationale":"Relevant corpora, tools, standards, and a 60-record retrospective design exist, but no opened evidence supplies the necessary fully coded authorized dataset, gold-standard material variants, reviewer availability, or preregistered tolerances.","source_ids":["S2","S3","S4","S5","S7"]},"safety_net_benefit":{"score":4,"rationale":"Untouched-source links, mandatory full display, random full-record audits, version checks, and immediate fallback directly address false normality and suppressed context. Their effectiveness still requires adversarial and subgroup testing.","source_ids":["S1","S4","S8"]},"scalability":{"score":2,"rationale":"Software components are reusable, but models, ontologies, permissions, sensitivity classes, translations, and materiality criteria must be rebuilt or reauthorized for each collection. Section-dependent Vedic results argue against broad transfer.","source_ids":["S2","S3","S7","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Resource-equivalent cost for the proposed 60-record retrospective experiment: permission and governance confirmation, transparent baseline model, three review conditions, audit sample, analysis, and preregistration.","confidence":"LOW","assumptions":["Records are already digitized, authorized, and fully coded.","Existing open-source collation and annotation components are reused.","Approximately 6–12 person-weeks are required across a domain scholar, research assistant, data analyst, and fractional developer.","No translation, new transcription, licensing purchase, or community travel is required.","The band is a bottom-up resource estimate, not a vendor quote."],"source_ids":["S3","S4","S5","S7"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Harden one-collection prototype into a governed institutional research service with authentication, provenance, audit logs, source viewer, model registry, accessibility testing, documentation, and security/privacy review.","confidence":"LOW","assumptions":["One institution and one coding scheme are in scope.","The service remains advisory and read-only.","Existing corpus infrastructure and TEI-compatible data are available.","Community consultation and governance work are included but no new digitization campaign is included."],"source_ids":["S2","S3","S4","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch across several authorized collections with collection-specific models, governance agreements, reviewer training, independent audits, monitoring, support, and integration with existing scholarly infrastructure.","confidence":"LOW","assumptions":["Three to five heterogeneous collections are included.","Each collection requires separate ontology validation, permission review, bypass policy, and held-out evaluation.","At least one full-time technical role plus substantial scholarly and stewardship time is required during launch.","No mass manuscript imaging, transcription, or translation program is included."],"source_ids":["S2","S3","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Annual resource equivalent for hosting, support, model and schema review, permission checks, periodic full-record audits, incident response, accessibility maintenance, and collection-specific recalibration.","confidence":"LOW","assumptions":["A small multi-collection service is maintained after launch.","One fractional engineer or research-software specialist and fractional scholarly, governance, and audit effort are retained.","Major corpus acquisition, digitization, and translation remain out of scope.","Audit frequency rises if reconstruction or subgroup errors approach stopping thresholds."],"source_ids":["S2","S3","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent academic projects document nonuniform ritual witnesses, consequential variants, repetitive material, and corpora whose volume and fragmentation require computational infrastructure. The specific attention-cost magnitude remains an empirical gap.","source_ids":["S1","S2","S6","S7"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"T-ReS, the Mainz Hittite Festive Rituals project, and the Trier liturgical-edition project are identifiable organizations already addressing closely related scholarly needs. This verifies credible potential adopters or authorizers, not commitment to this proposal.","source_ids":["S1","S2","S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The claim can be tested against full-record review and nonpredictive collation on time, material-variant recall, reconstruction disagreement, subgroup/context misses, bypass compliance, and total maintenance burden.","source_ids":["S4","S5","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A frozen 60-record, single-corpus, retrospective, read-only, preregistered comparison is bounded in records, scope, permissions, users, outcomes, and stopping rules.","source_ids":["S3","S4","S5"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step can be restricted to already authorized records, no substantive religious judgments, full source retention, mandatory bypasses, and collection-level shutdown. It must not start until corpus and any applicable community authority confirm the permitted uses.","source_ids":["S4","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four bands specify resource-equivalent scope and major exclusions and distinguish evidence generation, startup, launch, and recurrence. Confidence is low because no procurement quotes, local salary schedule, or measured coding effort was obtained.","source_ids":["S2","S3","S4","S5"]}},"next_evidence_step":"With a named corpus owner and any required community steward, preregister a read-only study of 60 already authorized, fully coded records from one collection. Fit and freeze a transparent sequence model on the earliest 30; reserve 30 untouched walk-forward records. Randomize and counterbalance qualified reviewers across three conditions: (A) complete-record sequential review, (B) nonpredictive TEI/CollateX-style apparatus or diff with full source access, and (C) the proposed predictive-residual interface with identical source access. Use an independent adjudication panel, blinded to condition, to define prespecified material variants and review every mandatory-bypass record plus a random and risk-stratified sample of collapsed records. Measure total person-time including setup and audit, material-variant recall, false discoveries, complete-sequence reconstruction disagreement, source expansions, coder confidence, errors by collection/context representation, bypass failures, fallback rate, and governance incidents. Falsify continuation if condition C misses any mandatory-bypass item, is more than 5 percentage points worse than A on material-variant recall or reconstruction agreement, shows a greater than 5-point disparity in miss rate for underrepresented contexts, fails to save at least 20 percent total person-time after audit and maintenance, or performs no better than B. The step authorizes only reject, revise, or conduct a larger test—not deployment or substantive religious conclusions.","blocking_evidence":["No direct observation establishes that full-record rereading is a binding attention bottleneck in the selected corpus.","No held-out result establishes reconstruction fidelity, material-variant recall, or noninferiority to full-record and nonpredictive-collation workflows.","No adopter has expressed demand for predictive residual review or committed personnel, corpus access, or funding.","No corpus-specific permission and community-authority determination has been documented for model training, residual display, audit logging, and sensitive outlier handling.","No validated ritual event-code ontology or evidence shows that it preserves performance, silence, translation ambiguity, material practice, and interpretively important routine wording.","No measured engineering, coding, governance, audit, or recurring support effort supports the dollar bands.","No evidence establishes whether the interface itself contaminates later coding or induces novelty bias."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The bounded search found established standards, products, institutional practices, and primary research for variant apparatuses, automated collation, religious-corpus analysis, repetition visualization, ritual text-reuse measurement, synoptic digital editions, and community data governance. It did not establish the complete predictive-residual review-and-fallback loop or its claimed net benefit. This is not a world-novelty, patentability, freedom-to-operate, market-size, or realized-impact determination.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain written pilot interest and authorization from one corpus-owning institution and all applicable community data authorities.","Measure baseline full-record review time and confirm that repeated expected units are a binding cost.","Freeze an authorized 60-record dataset, event-code ontology, material-variant adjudication protocol, and mandatory-bypass policy.","Complete the preregistered three-condition comparison and report total-time savings, recall, reconstruction, subgroup/context errors, and governance incidents.","Demonstrate zero mandatory-bypass misses and meet the predeclared noninferiority, equity, and net-time thresholds.","Replace resource assumptions with measured labor and infrastructure costs before any launch decision."],"reason":"Bounded web research verifies the domain problem, credible institutional actors, mature adjacent tooling, and a distinct testable joint claim, but cannot establish the claimed attention savings, fidelity, subgroup safety, workflow effects, adopter commitment, or real costs. Those questions require proprietary or governed corpus access, human reviewer fieldwork, and live comparative testing; under the controller rule this requires an empirical-research stop."},"proposal_index":1}