{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"invariant_mode_decomposition_design__literature_literary_theory:P1:v0","cell_id":"invariant_mode_decomposition_design__literature_literary_theory","search_queries":["site:cambridge.org unreliable narration reader judgments reliability ambiguity empirical study","computational analysis manuscript revisions version history literary editing stylometry PCA","site:the-efa.org developmental editing author editorial brief revisions","site:copyright.gov unpublished manuscript copyright editor data analysis","computational genetic criticism manuscript revision visualization versions literary drafts paper","digital scholarly editing version comparison literary drafts computational analysis genetic criticism","empirical study unreliable narrator reader response reliability judgments fiction","narrative unreliability computational detection unreliable narrator NLP paper","\"Understanding Iterative Revision from Human-Written Text\" ACL Anthology","site:tei-c.org release guidelines critical apparatus textual variation revisions manuscripts","site:nist.gov linear regression sample size overfitting validation model residuals guidance","site:statsmodels.org OLS multicollinearity small sample bootstrap regression","site:the-efa.org rates developmental editing fiction 2026 rate chart","university IRB reader study literary research consent guidance official minimal risk","site:hhs.gov human subjects research informed consent identifiable private information interviews surveys","site:ico.org.uk data protection research participant data literary manuscript consent","\"Trust in Stories\" \"In a Grove\" readers initially accept narrator revise judgment 2025","\"Trust in Stories: A Reader Response Study\" PDF","site:mdpi-res.com \"literature-05-00024\""],"sources":[{"source_id":"S1","title":"Trust in Stories: A Reader Response Study of (Un)Reliability in Akutagawa’s “In a Grove”","publisher":"MDPI, Literature","url":"https://www.mdpi.com/2410-9789/5/4/24","source_class":"PRIMARY_RESEARCH","publication_date":"2025-09-30","accessed_at":"2026-08-03","claims_supported":["An exploratory study used 148 readers and mixed quantitative and qualitative methods to study trust judgments in conflicting narrative accounts.","Narrative-trust judgments varied with sequencing, textual inconsistencies, character roles, and reader context.","Readers sometimes revised reliability judgments after encountering contradictions or paratextual cues.","The study itself describes empirical work on actual readers' negotiation of narrative reliability as sparse."]},{"source_id":"S2","title":"Classifying Unreliable Narrators with Large Language Models","publisher":"Association for Computational Linguistics","url":"https://aclanthology.org/2025.acl-long.1013/","source_class":"PRIMARY_RESEARCH","publication_date":"2025-07","accessed_at":"2026-08-03","claims_supported":["TUNa is a human-annotated dataset covering multiple forms of narratorial unreliability and including literary narratives.","Computational classification of unreliable narrators is established adjacent research.","The authors report that unreliable-narrator classification remains challenging."]},{"source_id":"S3","title":"Understanding Iterative Revision from Human-Written Text","publisher":"Association for Computational Linguistics","url":"https://aclanthology.org/2022.acl-long.250/","source_class":"PRIMARY_RESEARCH","publication_date":"2022-05","accessed_at":"2026-08-03","claims_supported":["Writing revision is iterative and can occur at multiple depths and granularities.","IteraTeR operationalizes iterative revision with edit-intention annotations across domains.","The paper found substantial ambiguity between clarity, coherence, and style annotations and poor correspondence between some automatic quality metrics and manual evaluation.","Computational modeling of revision trajectories is established prior research, although not for unreliable-narrator manuscript editing."]},{"source_id":"S4","title":"A Computational Approach to Walt Whitman's Stylistic Changes in Leaves of Grass","publisher":"arXiv","url":"https://arxiv.org/abs/2111.05414","source_class":"PRIMARY_RESEARCH","publication_date":"2021-11-09","accessed_at":"2026-08-03","claims_supported":["Computational comparison across seven editions of one literary work is established adjacent practice.","The study used PCA on textual features to visualize stylistic change across editions.","Its PCA is descriptive dimensionality reduction, not a fitted transition operator, modal stability test, reader-linked checkpoint, or prospective edit intervention.","The source is a preprint and therefore provides weaker validation than peer-reviewed field evidence."]},{"source_id":"S5","title":"TEI P5 Guidelines, Chapter 13: Critical Apparatus","publisher":"Text Encoding Initiative Consortium","url":"https://www.tei-c.org/release/doc/tei-p5-doc/en/html/TC.html","source_class":"STANDARD","publication_date":"2026-02-18","accessed_at":"2026-08-03","claims_supported":["A mature standard exists for encoding variant readings across manuscripts and editions.","The standard supports structured apparatus entries and interactive selection among witness readings.","Editorial methodology affects how variation is identified and represented.","TEI represents textual variation but does not supply the proposed transition-matrix, stability, reader-outcome, or edit-sensitivity analysis."]},{"source_id":"S6","title":"Editorial Service Definitions","publisher":"Editorial Freelancers Association","url":"https://www.the-efa.org/editorial-services-definitions/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-08-11","accessed_at":"2026-08-03","claims_supported":["Developmental editors address content, organization, and genre and commonly deliver manuscript evaluations or revision letters.","Line editing acts at sentence or paragraph level and focuses on language and style.","The EFA reports more than 4,000 members, establishing an identifiable practitioner class.","The page does not express demand for modal analysis or document the proposed reliability-amplification failure."]},{"source_id":"S7","title":"2026 Editorial Rates","publisher":"Editorial Freelancers Association","url":"https://www.the-efa.org/wp-content/uploads/2026/03/2026-Rate-Chart.pdf","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2026-03","accessed_at":"2026-08-03","claims_supported":["Reported 2026 fiction developmental-editing rates are approximately $52.50–$70 per hour.","Reported fiction line-editing rates are approximately $50–$60 per hour.","Reported beta-reading rates are approximately $50–$70 per hour, while manuscript assessment has a reported $600 median project price.","These rates support only editorial and reader labor assumptions, not specialized statistical analysis, software, recruitment, governance, or institutional-review costs."]},{"source_id":"S8","title":"How can I tell if a model fits my data?","publisher":"National Institute of Standards and Technology","url":"https://www.itl.nist.gov/div898/handbook/pmd/section4/pmd44.htm","source_class":"OFFICIAL_GUIDANCE","publication_date":"NOT_STATED","accessed_at":"2026-08-03","claims_supported":["Residual analysis is a primary method for evaluating fitted-model adequacy.","Residuals are observed responses minus model predictions.","Validation becomes difficult when the number of estimated parameters is close to the data-set size.","The guidance supports residual and holdout scrutiny but does not validate literary interpretation of fitted modes."]}],"problem_evidence":{"support":"WEAK","rationale":"The sources establish that revision is iterative, that clarity/coherence/style labels can be difficult to separate, and that reader judgments of narratorial reliability are dynamic and context-sensitive. No source directly documents developmental editors unintentionally amplifying a stable coupled direction that moves a novel outside an author's intended ambiguity envelope. The stated failure is plausible but not visibly demonstrated.","source_ids":["S1","S3","S6"]},"stakeholder_evidence":{"support":"WEAK","rationale":"The EFA establishes a large, identifiable developmental-editing practitioner class and ordinary manuscript-evaluation workflow. No named author, editor, publisher, funder, or professional organization expresses need for a modal ambiguity checkpoint or willingness to adopt or authorize one. The author-as-authorizer exists only in the proposal, not in external evidence.","source_ids":["S6","S7"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Ordinary developmental and line editing","similarity":"Evaluates manuscript-level genre, organization, language, and style and proposes revisions to an author.","remaining_difference":"It does not estimate a repeated multivariate transformation, modal gains, held-out residuals, or edit-level modal sensitivity.","source_ids":["S6"]},{"name":"IteraTeR iterative-revision modeling","similarity":"Represents successive revisions computationally and distinguishes revision intentions, depths, and granularities.","remaining_difference":"It addresses general formal writing and revision quality, not manuscript-specific invariant directions tied to reader judgments and author-declared ambiguity.","source_ids":["S3"]},{"name":"TUNa unreliable-narrator classification","similarity":"Operationalizes narratorial unreliability with human annotations and computational models.","remaining_difference":"It classifies narrative accounts; it does not model transitions between manuscript drafts or recommend reversible editorial batches.","source_ids":["S2"]},{"name":"Empirical reader-response study of narrative trust","similarity":"Measures how actual readers assign and revise reliability judgments under conflicting accounts and sequencing changes.","remaining_difference":"It tests reception of a fixed story rather than the effect of successive editorial transformations or alternative edit packets.","source_ids":["S1"]},{"name":"Computational analysis of Whitman editions using PCA","similarity":"Uses multivariate textual features and PCA to characterize change across versions of a literary work.","remaining_difference":"It is retrospective descriptive dimensionality reduction, not a locally fitted revision operator with stability, intervention, residual, and reader-outcome tests.","source_ids":["S4"]},{"name":"TEI critical apparatus","similarity":"Provides structured representation of variant readings across manuscript or edition witnesses.","remaining_difference":"It records and exposes variants but does not estimate coupled revision modes, gains, reader associations, or prospective edit leverage.","source_ids":["S5"]}],"distinctive_claim_remaining":"For one manuscript and a declared revision phase, a resampling-stable coupled direction fitted from earlier revision transitions will improve held-out prediction of textual change and blinded-reader reliability outcomes over the best coordinate-wise, checklist, and whole-version baselines; projecting proposed edits onto that direction will also identify a reversible alternative batch that better stays within the author's ambiguity envelope without materially reducing event reconstructability.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Variant encoding, human annotation, reader-response measurement, computational revision modeling, PCA, and residual validation are all technically precedented. A six-coordinate transition model is computationally routine. Feasibility is nevertheless constrained by subjective and potentially theory-laden coding, manuscript confidentiality, dependence among passages, uncertain stationarity across editorial rounds, possible non-normal or weak-gap operators, an underspecified reader sample, and the limited effective sample behind the proposed 30-passage holdout. No source validates eigenmode interpretation for literary revision or establishes the required legal, institutional-review, recruitment, or data-governance workflow.","source_ids":["S1","S2","S3","S4","S5","S8"]},"scores":{"meaningful_impact":{"score":2,"rationale":"Preserving intended ambiguity could matter artistically, but neither prevalence nor consequential harm from the specific coupled failure is externally demonstrated; realized impact is unmeasured.","source_ids":["S1","S3"]},"stakeholder_pull":{"score":1,"rationale":"A relevant practitioner class exists, but no specific adopter, authorizer, or funder has expressed demand for this checkpoint.","source_ids":["S6","S7"]},"incremental_advantage":{"score":2,"rationale":"The proposed checkpoint could expose joint movements missed by coordinate reviews, but no evidence shows better prediction or editorial decisions than close reading, checklists, or blinded version comparison.","source_ids":["S1","S3","S4","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The searched analogues separately cover revision modeling, unreliable-narrator classification, reader studies, PCA across editions, and variant apparatuses; none located combines a fitted revision operator, stability testing, reader linkage, and reversible edit sensitivity. This is not a world-novelty finding.","source_ids":["S1","S2","S3","S4","S5"]},"technical_implementability":{"score":2,"rationale":"The component methods are implementable, but stable manuscript-specific eigenmodes require more evidence on sample adequacy, coder reliability, passage dependence, stationarity, conditioning, and out-of-sample validity.","source_ids":["S3","S4","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"An author can authorize a read-only retrospective study and retain final artistic control, making authority conceptually straightforward. Feasibility remains contingent on locating a willing rights holder, editor, and reader-study host and executing confidentiality and research-review arrangements.","source_ids":["S6"]},"evidence_readiness":{"score":2,"rationale":"The proposal has measurable coordinates, comparators, holdouts, and falsifiers, but lacks a partner, proprietary snapshots, powered reader sample, validated coding instrument, and direct evidence of the problem.","source_ids":["S1","S3","S8"]},"safety_net_benefit":{"score":2,"rationale":"Reversibility, residual checks, and author veto reduce risk of automatic artistic prescription, but the checkpoint may still reify unstable mathematical modes, expose unpublished text, or displace unmeasured aesthetic qualities.","source_ids":["S1","S8"]},"scalability":{"score":1,"rationale":"Each deployment depends on a manuscript-specific authorial brief, multiple snapshots, trained coding, reader testing, and refitting after phase changes; external validity across manuscripts is expressly unsupported.","source_ids":["S1","S3","S4"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One retrospective, read-only study of one manuscript: protocol and power analysis, 60 matched passages across four snapshots, two coders plus adjudication, locked model fitting, approximately 120 blinded reader sessions, comparator analysis, and a nonbinding report.","confidence":"MODERATE","assumptions":["Existing manuscript snapshots and edit logs are supplied without acquisition cost.","Coder/editor labor is valued near the EFA's reported fiction-editing range of roughly $50–$70 per hour.","Reader labor is valued near the EFA beta-reading range before recruitment overhead.","Specialist statistical labor, recruitment, secure storage, and research-review administration are assumed at higher uncited rates.","If power analysis requires substantially more than 120 readers, this band may be exceeded."],"source_ids":["S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Convert a successful study into a controlled service: validated coding manual and software, secure manuscript/version pipeline, reproducible analysis, audit/report templates, staff training, confidentiality controls, and governance thresholds.","confidence":"LOW","assumptions":["No automatic rewriting or manuscript decision system is built.","One organization supports a small analyst/editor team.","Bespoke security, legal review, and method validation dominate costs.","No direct market quote was found for this specialized literary-statistical workflow."],"source_ids":["S5","S7","S8"]},"operational_launch":{"band_2026_usd":"10K_TO_50K","scope":"Run the first prospective checkpoint for one consented manuscript phase, including recoding, reader measurement, comparison of original and simulated alternative packets, author/editor review, and rollback-ready reporting.","confidence":"LOW","assumptions":["The startup tooling already exists.","The run covers one substantial edit batch rather than a full novel lifecycle.","Reader and coder effort is similar to the first-evidence study.","A major rewrite or failed drift threshold triggers refitting and may move the work into a higher band."],"source_ids":["S6","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain a limited program supporting roughly four to eight manuscript checkpoints per year, including analyst time, coders/readers, secure storage, drift checks, quality review, and method maintenance.","confidence":"LOW","assumptions":["Volume remains low and bespoke.","Each manuscript requires separate authorization and local fitting.","No cross-author training corpus is created.","Specialized analysis, recruitment, governance, and security costs are not priced in the available editorial rate data."],"source_ids":["S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"NO","reason":"External evidence supports iterative revision and context-sensitive reader reliability judgments, but not the specific recurring reliability-amplification failure during novel editing.","source_ids":["S1","S3","S6"]},"externally_credible_adopter_or_authorizer":{"status":"NO","reason":"The EFA establishes a relevant editor population, but no named author, editor, publisher, institution, or funder expresses need or offers authorization for this work.","source_ids":["S6","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim compares locked modal predictions and alternative edit packets against coordinate-wise regression, an editorial checklist, and blinded whole-version ratings on held-out data, with clear failure conditions.","source_ids":["S1","S3","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-manuscript retrospective study can be bounded by fixed snapshots, passages, coders, reader sessions, comparators, preregistered thresholds, and a prohibition on changing the working manuscript.","source_ids":["S1","S3","S7","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"Author veto and read-only analysis are protective, but no actual rights holder, confidentiality agreement, reader-research determination, data-retention plan, or institutional host has been secured.","source_ids":["S1","S6"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Current EFA rates support rough editorial and reader labor estimates, but specialized statistical, recruitment, security, legal, and governance costs lack direct pricing evidence.","source_ids":["S7"]}},"next_evidence_step":"Secure one consenting author-editor pair with four already-existing manuscript snapshots and documented revision goals. Before accessing text, obtain the host institution's human-participant determination and execute manuscript-confidentiality and data-destruction terms. Preregister six coordinates, reader outcomes, passage-clustering treatment, coder agreement threshold, residual tolerance, resampling stability, spectral-gap rule, and a power analysis. Score 60 matched passages with two blinded coders and adjudicate disagreements; fit and lock the transition model on the first two transitions and test the third. Recruit approximately 120 blinded readers only if the power analysis supports that bound. Compare against coordinate-wise regularized regression, the existing editorial checklist, blinded whole-version ratings, and a no-change/null model. Require at least a preregistered practically meaningful held-out error improvement over the best rival, stable mode direction under passage and coder resampling, and a simulated alternative packet that improves ambiguity-envelope compliance while meeting a noninferiority margin for event reconstructability. Falsify the intervention if no stable coupled direction survives resampling, the best coordinate rival matches or beats prediction, reader associations reverse across subsets, or the alternative packet provides no advantage. Do not modify the working manuscript.","blocking_evidence":["No direct prevalence evidence that ordinary editing repeatedly amplifies narratorial reliability through a stable coupled direction.","No named author, editor, publisher, institution, or funder has expressed need or granted access.","No manuscript snapshots, edit logs, or documented authorial ambiguity envelope are available for testing.","No validated coding instrument or demonstrated inter-coder reliability exists for the six proposed coordinates.","No power analysis establishes that the proposed passage and reader samples can distinguish a stable transition mode from noise and passage dependence.","No held-out comparison shows predictive advantage over coordinate-wise regression, editorial checklists, or blinded version ratings.","No prospective evidence shows that modal projection yields a better reversible edit packet.","Human-participant review, manuscript-rights terms, confidentiality controls, and data-retention authority remain unresolved.","Specialized analysis, governance, recruitment, and security costs are not directly priced."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search located only adjacent practices and research, not an exact combination. This bounded review does not measure world novelty, patentability, freedom to operate, market size, realized impact, or exhaustive prior art; absence of a close match in eight sources is not evidence of novelty.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Obtain a named author-editor partner, documented artistic brief, and authorized access to four revision snapshots.","Demonstrate acceptable coder reliability for the six textual coordinates using a preregistered instrument.","Establish sample adequacy through power analysis and passage-cluster-aware validation.","Show a resampling-stable coupled mode on training transitions that remains identifiable on the held-out transition.","Beat coordinate-wise regression, the editorial checklist, blinded whole-version ratings, and a null model by preregistered practical margins.","Show that a reversible simulated alternative packet better preserves the ambiguity envelope without materially reducing reconstructability.","Resolve human-participant review, manuscript confidentiality, data retention, and rights-holder authority.","Replace low-confidence resource assumptions with partner-specific staffing, recruitment, security, and governance quotes."],"reason":"Bounded web research establishes adjacent methods but cannot verify the manuscript-specific failure, adopter demand, stable mode, predictive advantage, or actionable edit benefit. Those questions require proprietary revision histories, trained coding, blinded readers, and live or retrospective field testing. Under the required controller rule, this empirical stop is not repairable within web research and therefore sets repairable false."},"proposal_index":1}