{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"invariant_mode_decomposition_design__literature_literary_theory:P4:v0","cell_id":"invariant_mode_decomposition_design__literature_literary_theory","search_queries":["digital humanities genre classification corpus machine learning instability genre labels iterative clustering prototype classification","computational literary genre classification hybrid genres limitations corpus metadata","iterative classification self-training feedback loops label instability prototypes research","digital library genre metadata controlled vocabulary curator expressed need reproducibility","site:methods.clsinfra.io corpus building genre metadata quality multiple genre labels literary","performative prediction repeated retraining feedback stable point paper PMLR","self-training confirmation bias pseudo-labels iterative classification paper primary research","prototype-based classification iterative prototype update labels stability paper","site:bls.gov oes data scientists librarians curators hourly wage May 2025","site:living-with-machines.github.io genre classification ambiguous multiple labels crowdsourcing workflow","site:scikit-learn.org stable LabelSpreading semi-supervised documentation iterative labels","genre classification computational literary studies corpus reproducibility metadata labels primary study"],"sources":[{"source_id":"S1","title":"Evaluation in Genre Analysis","publisher":"CLS INFRA","url":"https://methods.clsinfra.io/evaluation-genre.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2023","accessed_at":"2026-08-03","claims_supported":["Hybrid, partial, and multiple genre assignments complicate evaluation, yet are commonly reduced to categorical classes.","Genre-classification evaluation depends on corpus design and should compare multiple approaches.","Hybrid-genre evaluation remains an open problem; probabilistic multi-label and hierarchical classification are recognized responses."]},{"source_id":"S2","title":"Corpus Building for Genre Analysis","publisher":"CLS INFRA","url":"https://methods.clsinfra.io/corpus-genre.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2023","accessed_at":"2026-08-03","claims_supported":["Genre categories are difficult to define coherently, are tradition-dependent, and individual texts can participate in multiple genres.","Genre metadata materially determines corpus composition and downstream analysis.","Copyrighted literary corpora can require jurisdiction- and project-specific text-and-data-mining authorization or derived-text access controls."]},{"source_id":"S3","title":"MODS User Guidelines: genre","publisher":"Library of Congress","url":"https://www.loc.gov/standards/mods/v3/genre","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Genre is repeatable metadata and may use controlled vocabularies with explicit authorities.","Institutions should apply genre terms consistently and document their approach.","Genre metadata supports indexing, filtering, and research discovery, making release consistency consequential."]},{"source_id":"S4","title":"Genre Classification — Classifying 19th Century British Library Books Using Crowdsourcing and Machine Learning","publisher":"Living with Machines consortium and British Library","url":"https://living-with-machines.github.io/genre-classification/genre_classification.html","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2021","accessed_at":"2026-08-03","claims_supported":["An identifiable GLAM consortium explored machine-learning genre classification for British Library books.","GLAM institutions express growing interest in machine-assisted metadata creation and evaluation at scales difficult for manual processing.","The project explicitly characterizes genre as complex, contested, and temporally changeable and calls its own classification crude."]},{"source_id":"S5","title":"Performative Prediction","publisher":"Proceedings of Machine Learning Research","url":"https://proceedings.mlr.press/v119/perdomo20a.html","source_class":"PRIMARY_RESEARCH","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["Predictions or decisions can influence later target data, producing feedback-driven distribution shift.","Repeated retraining has a formal stability problem, with convergence depending on sensitivity conditions.","Performative stability is established adjacent prior art for studying model-data feedback."]},{"source_id":"S6","title":"Debiased Self-Training for Semi-Supervised Learning","publisher":"arXiv","url":"https://arxiv.org/abs/2202.07136","source_class":"PRIMARY_RESEARCH","publication_date":"2022-02-15","accessed_at":"2026-08-03","claims_supported":["Iterative pseudo-label reuse can accumulate errors, amplify category imbalance, and create training instability.","Anchoring label generation in clean labeled cases and decoupling it from pseudo-label utilization are tested controls against confirmation bias.","The reported experiments are on machine-learning benchmarks, not literary-corpus genre governance."]},{"source_id":"S7","title":"LabelSpreading — scikit-learn 1.9.0 Documentation","publisher":"scikit-learn developers","url":"https://scikit-learn.org/stable/modules/generated/sklearn.semi_supervised.LabelSpreading.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["Production-grade open-source tooling already supports iterative label spreading with soft clamping, convergence tolerances, probabilistic label distributions, and iteration limits.","Soft retention of initial labels and steady-state checks are established implementable controls adjacent to the proposed anchors and gates."]},{"source_id":"S8","title":"Interpretable Outputs: Criteria for Machine Learning in the Humanities","publisher":"Digital Humanities Quarterly","url":"https://dhq.digitalhumanities.org/vol/15/2/000555/000555.html","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2021","accessed_at":"2026-08-03","claims_supported":["Humanities classification requires access to features, weights, parameters, and interpretable evidence rather than aggregate accuracy alone.","Different algorithms can produce similar accuracy while relying on substantially different features.","Computational humanities workflows must preserve ambiguity and ground claims in objects interpretable by the relevant scholarly community."]}],"problem_evidence":{"support":"MODERATE","rationale":"The constituent problem is visible and consequential: hybrid or multi-genre works challenge categorical evaluation, genre metadata shapes research cohorts, and iterative self-training can amplify errors and category imbalance. However, no opened source documents a literary corpus that currently recomputes genre prototypes from its own assignments or demonstrates the proposed oscillatory coupled reassignment mode. The proposal therefore extrapolates from two supported problems—genre ambiguity and algorithmic feedback—rather than verifying their conjunction in an operating corpus.","source_ids":["S1","S2","S3","S5","S6"]},"stakeholder_evidence":{"support":"WEAK","rationale":"The Library of Congress states that consistent genre metadata supports filtering and research, while the Living with Machines/British Library collaboration demonstrates identifiable GLAM interest in machine-assisted genre metadata. These establish plausible curator and institutional stakeholders, but none expresses a need for Jacobian decomposition, a spectral stability gate, or even an assignment-dependent prototype-recalibration workflow. No current corpus owner, authorized slice, funder commitment, or named decision maker for a pilot was found.","source_ids":["S3","S4"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Performative prediction and repeated retraining stability","similarity":"Formalizes feedback in which model outputs affect later data and analyzes whether repeated retraining converges to stable points.","remaining_difference":"It does not address literary genre metadata, multi-label scholarly adjudication, local spectral modes of work–genre memberships, or a release gate.","source_ids":["S5"]},{"name":"Debiased self-training with clean-label anchoring","similarity":"Directly addresses iterative pseudo-label error accumulation and separates clean-label anchors from feedback-prone pseudo-label use.","remaining_difference":"Its controls and validation concern generic benchmark learning performance rather than curator-governed metadata releases, modal gains, cohort reproducibility, or preservation of disputed literary memberships.","source_ids":["S6"]},{"name":"LabelSpreading with soft clamping and convergence tolerance","similarity":"Provides an established iterative soft-label method with explicit anchoring strength, probabilistic assignments, iteration limits, and convergence criteria.","remaining_difference":"It does not diagnose a curator's arbitrary prototype-update Jacobian, compare coupled unstable modes, or trigger scholarly review based on held-out corpus outcomes.","source_ids":["S7"]},{"name":"Living with Machines British Library genre-classification workflow","similarity":"Applies crowdsourcing and machine learning to literary genre metadata in an identifiable GLAM setting while acknowledging genre complexity.","remaining_difference":"The documented workflow does not derive prototypes from current assignments, repeatedly recalibrate memberships, analyze spectral stability, or gate prototype contributions.","source_ids":["S4"]}],"distinctive_claim_remaining":"For an authorized literary corpus that actually uses assignment-dependent prototype recalibration, a resampling-stable local modal model plus sensitivity-selected one-cycle contribution gate will predict and reduce held-out unsupported membership switching and cohort divergence better than per-work confidence review, aggregate-count monitoring, fixed anchors alone, and non-transformational clustering, while preserving all disputed multi-label memberships and not reducing agreement with independently adjudicated reference cases.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"The numerical components are implementable on a 150-work, three-genre slice using ordinary finite-difference Jacobian estimation, eigendecomposition, resampling, and held-out replay; adjacent software already implements iterative probabilistic labels, soft clamping, and convergence controls. Workflow feasibility is conditional on a fully documented deterministic recalibration function, versioned starting states, traceable evidence, and curator-controlled sandbox access. Scholarly adjudication is indispensable because modes cannot define genre meaning. Copyright or access authorization is corpus- and jurisdiction-specific. Primary safety risks are epistemic: freezing inherited bias, suppressing hybrid traditions, or laundering a local numerical result into an essentialist classification; read-only execution, preserved multi-label metadata, interpretability reporting, and curator rollback mitigate but do not validate these risks.","source_ids":["S2","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Stable genre metadata can materially affect corpus retrieval, composition, and reproducibility, but the prevalence and magnitude of feedback-driven release divergence are unmeasured.","source_ids":["S2","S3"]},"stakeholder_pull":{"score":2,"rationale":"GLAM organizations show interest in machine-assisted metadata, but no adopter has requested or authorized this specific diagnostic or gate.","source_ids":["S3","S4"]},"incremental_advantage":{"score":3,"rationale":"Coupled held-out trajectory prediction and intervention sensitivity could reveal failures missed by isolated confidence scores, but superiority has not been tested against simpler anchors, soft clamping, or adjudication.","source_ids":["S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"No direct literary-corpus spectral release gate was found, while core ingredients—feedback stability, clean anchors, clamping, and iterative convergence checks—are established adjacent art.","source_ids":["S4","S5","S6","S7"]},"technical_implementability":{"score":3,"rationale":"The bounded computation is conventional if an explicit, repeatable calibration operator exists; numerical fragility, nonlinearity, and absent trajectories could invalidate the local model.","source_ids":["S5","S7"]},"adoption_authority_feasibility":{"score":2,"rationale":"Corpus curators plausibly control metadata releases, but no actual curator, governance process, licensing basis, or authorized corpus slice is identified.","source_ids":["S2","S3","S4"]},"evidence_readiness":{"score":2,"rationale":"External sources support genre ambiguity and generic feedback instability, but not the candidate's exact problem prevalence, target workflow, or outcome advantage.","source_ids":["S1","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"A nonpublishing sandbox, unchanged public snapshot, preserved disputed labels, explicit interpretation limits, and rollback to manual review create a strong reversible evidence path.","source_ids":["S3","S8"]},"scalability":{"score":3,"rationale":"Computation should scale for sparse work–genre states, but scholarly reference adjudication, rights review, and repeated mode validation remain human bottlenecks.","source_ids":["S2","S4","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One preregistered sandbox study on at most 150 works, three genres, five replay rounds, four comparators, and a small independently adjudicated multi-label reference set.","confidence":"LOW","assumptions":["A documented recalibration rule, starting snapshot, and textual-evidence representation already exist.","Approximately 160–400 combined hours from a technical analyst, curator/corpus engineer, and genre scholars.","Existing institutional compute and open-source numerical tooling are used.","This is a resource-equivalent estimate, not a vendor quote; the wage search did not yield a relied-upon direct cost source."],"source_ids":["S7"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Productionize reproducible replay, versioned state extraction, access controls, audit reports, monitoring thresholds, curator review screens, and rollback integration for one corpus.","confidence":"LOW","assumptions":["One existing metadata platform and release pipeline are modified rather than replaced.","Roughly 0.5–1.5 staff-years across engineering, data analysis, curatorial governance, and security/rights review.","No full-text licensing acquisition or taxonomy redesign is included."],"source_ids":["S2","S3","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Validate across several corpus slices and release cycles, complete scholar adjudication and acceptance testing, train operators, and authorize the first governed production release.","confidence":"LOW","assumptions":["Launch includes multiple preregistered replications and failure drills.","Disputed memberships remain visible; no mass manual recataloguing is included.","Institutional legal, infrastructure, and curator functions already exist."],"source_ids":["S2","S3","S4"]},"annual_recurring":{"band_2026_usd":"10K_TO_50K","scope":"Per-release replay and monitoring for one bounded corpus, periodic re-estimation, report review, exception adjudication, software maintenance, and annual governance review.","confidence":"LOW","assumptions":["Two to four bounded releases or recalibrations occur annually.","Most runs are automated and only residual or gated cases receive manual review.","A persistent large disputed queue or taxonomy revision would raise this to the next band."],"source_ids":["S3","S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Genre ambiguity and iterative classification feedback are independently supported, but no source verifies that an operating literary corpus performs the stated self-referential prototype recalibration or suffers coupled oscillatory releases.","source_ids":["S1","S2","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"UNCERTAIN","reason":"The British Library/Living with Machines work and Library of Congress guidance identify plausible institutional actors and metadata needs, but there is no named current adopter, authorized slice, or expressed demand for this intervention.","source_ids":["S3","S4"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The claim specifies held-out trajectory and cohort outcomes, four comparators, preservation constraints, and reference-case noninferiority, so it can be falsified.","source_ids":["S1","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A 150-work, three-genre, five-round sandbox pilot is bounded and includes no gate, confidence review, fixed anchors, and sensitivity-selected modal gate comparators with explicit failure conditions.","source_ids":["S1","S7","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"The read-only design and rollback are sound, but actual curator authorization, corpus rights, reference adjudication, representational equity checks, and permission to process textual evidence remain unresolved.","source_ids":["S2","S3","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four resource stages are bounded with explicit staffing, infrastructure, scale, and exclusion assumptions, although confidence is low because no direct wage or vendor evidence was used.","source_ids":["S2","S3","S7","S8"]}},"next_evidence_step":"First secure a named corpus partner and verify that its documented rule truly recomputes prototypes from current assignments. On one authorized slice of no more than 150 works and three adjacent genres, preregister the state representation, perturbation radius, Jacobian estimator, resampling-stability and spectral-gap criteria, residual tolerance, cohort-overlap metric, unsupported-switching metric, reference-case noninferiority margin, subgroup/hybrid-tradition checks, and stopping rules. Replay five rounds from an immutable snapshot; fit only on early rounds and evaluate later rounds plus independently adjudicated multi-label cases. Randomize or counterbalance four sandbox arms: unchanged recalibration, per-work confidence review, fixed clean anchors/soft clamping, and the sensitivity-selected one-cycle modal gate. Falsify the proposal if no reproducible non-damped coupled mode appears; if the modal model does not improve held-out prediction over confidence and clustering baselines; if the gate does not reduce switching or divergence beyond fixed anchors; if reference-case agreement or hybrid-tradition representation worsens beyond the preregistered margin; or if mode identity, residuals, rights, or authority checks fail. Publish nothing and retain every candidate membership in the audit output.","blocking_evidence":["No externally documented literary corpus using assignment-dependent genre-prototype recalibration was found.","No named curator or institution has authorized a pilot or expressed demand for a spectral stability gate.","No versioned repeated-recalibration trajectories or prevalence estimate establish that coupled oscillation occurs in practice.","No independently adjudicated multi-label reference set, evidence rubric, or acceptable noninferiority margin is identified.","The local Jacobian's stationarity, numerical conditioning, spectral separation, and predictive validity are untested.","Corpus-specific copyright, data-access, and processing authority are unresolved.","The relative effect on hybrid or minor literary traditions requires live scholarly adjudication and cannot be established by further bounded web research.","Cost estimates lack direct wage, procurement, or institution-specific engineering evidence."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This review measured neither world novelty nor patentability, freedom to operate, market size, or realized impact. It found no direct match among the eight opened sources for a Jacobian/spectral stability release gate in literary genre metadata, but the search was bounded and did find established adjacent practices for feedback stability, self-training debiasing, clean anchoring, soft label clamping, convergence checks, multi-label genre representation, and interpretable humanities classification. Any novelty claim is therefore limited to the unvalidated corpus-governance combination, not its mathematical or workflow ingredients.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Obtain written participation and sandbox authorization from a named corpus curator whose live workflow demonstrably recomputes genre prototypes from current assignments.","Produce an immutable starting snapshot, five replayable calibration rounds, traceable evidence links, and an independently adjudicated multi-label reference set.","Preregister modal-conditioning, stability, residual, cohort, noninferiority, representation, rights, and rollback thresholds before analysis.","Run the four-arm held-out comparison and report all null results, mode-instability failures, residual cases, and subgroup effects.","Replace resource-equivalent cost assumptions with partner-specific labor, infrastructure, rights-review, integration, and recurring-operation estimates."],"reason":"The exact operating problem, adopter authority, and incremental effect cannot be resolved by additional bounded web search. They require proprietary or partner-controlled workflow data, live scholarly adjudication, authorization, and a sandbox experiment. Because every STOP must be non-repairable under the controller rule, repairable is false even though a future partnered study could generate a new candidate version."},"proposal_index":4}