{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"invariant_mode_decomposition_design__information_theory:P4:v0","cell_id":"invariant_mode_decomposition_design__information_theory","search_queries":["site:faa.gov readback hearback errors voice communication safety command confusion","site:eurocontrol.int air ground communication safety similar sounding words readback hearback","speech command vocabulary design confusion matrix acoustic confusability research","repeated relay spoken message errors Markov confusion matrix eigenvalue","iterated transmission spoken messages relay chain errors research confusion matrix","command vocabulary redesign replace confusable words speech recognition primary study","NASA voice command vocabulary confusability matrix confirmation dialogue design","official standard voice command vocabulary phonetic alphabet radio safety readback","eigenvalue analysis confusion matrix repeated communication relay speech commands","spectral decomposition stochastic confusion matrix command vocabulary redesign","Markov chain repeated human transmission message confusion experiment","\"Passing crisis and emergency risk communications\" Edworthy full text"],"sources":[{"source_id":"S1","title":"The Outcome of ATC Message Complexity on Pilot Readback Performance","publisher":"U.S. Federal Aviation Administration, Civil Aerospace Medical Institute","url":"https://www.faa.gov/data_research/research/med_humanfacs/oamtechreports/2000s/2006/200625","source_class":"PRIMARY_RESEARCH","publication_date":"2006-11","accessed_at":"2026-08-03","claims_supported":["A content analysis used 50 hours of operational pilot-controller communications from five busy terminal facilities.","Readback errors and pilot requests increased with message complexity and number of aviation topics.","Nonstandard phraseology produced observed communication problems and misunderstandings."]},{"source_id":"S2","title":"European Action Plan for Air Ground Communications Safety","publisher":"EUROCONTROL","url":"https://www.eurocontrol.int/publication/european-action-plan-air-ground-communications-safety","source_class":"OFFICIAL_GUIDANCE","publication_date":"2006-05-01","accessed_at":"2026-08-03","claims_supported":["Air-ground communication problems can create hazardous situations.","EUROCONTROL identified call-sign confusion, radio interference, simultaneous transmissions, and nonstandard phraseology as safety concerns.","The initiative collected occurrence reports and surveyed pilots and controllers, and its action plan called for implementation of recommendations."]},{"source_id":"S3","title":"Passing crisis and emergency risk communications: the effects of communication channel, information type, and repetition","publisher":"Applied Ergonomics (Elsevier); University of Plymouth research repository","url":"https://researchportal.plymouth.ac.uk/en/publications/passing-crisis-and-emergency-risk-communications-the-effects-of-c/","source_class":"PRIMARY_RESEARCH","publication_date":"2015-05","accessed_at":"2026-08-03","claims_supported":["Three experiments studied messages passed through chains of participants.","Verbal transmission was less accurate than written transmission in the first experiment.","Different message elements persisted differently along the chain, directly supporting the possibility of structured rather than uniform relay degradation.","Repetition improved spoken-message transmission, providing a relevant confirmation/repetition comparator."]},{"source_id":"S4","title":"Considerations for Implementing Voice-Controlled Spacecraft Systems through a Human-Centered Design Approach","publisher":"NASA Johnson Space Center","url":"https://ntrs.nasa.gov/api/citations/20180006618/downloads/20180006618.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"2018-08-01","accessed_at":"2026-08-03","claims_supported":["NASA defines and recommends vocabulary confusability-matrix testing for command-and-control speech applications.","The guidance addresses vocabulary selection, feedback, recognition-error handling, dialogue design, and usability testing.","NASA reports flight experience in which ground-trained command recognition declined in operation, while updated templates and confirmation dialogue improved performance.","The guidance states that task-level command accuracy can matter more than aggregate recognition accuracy.","It recommends replacing problematic vocabulary, engaging the recognizer vendor, evaluating operational noise and speech variability, and retaining alternative control when reliability degrades."]},{"source_id":"S5","title":"NASA-STD-3001 Volume 2, Section 10.0 Crew Interfaces","publisher":"National Aeronautics and Space Administration","url":"https://www.nasa.gov/reference/10-0-crew-interfaces-vol-2/","source_class":"STANDARD","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["NASA requires confirmation before completing critical, hazardous, irreversible, or destructive commands.","Operational nomenclature must be standardized and unambiguous.","Critical voice communications have an intelligibility requirement, while acknowledgement feedback and operating conditions must be considered.","Human operators must retain override, failure-recovery, and safe-control capabilities."]},{"source_id":"S6","title":"FAA Flight Services, Chapter 11: Phraseology, Section 1: General","publisher":"U.S. Federal Aviation Administration","url":"https://www.faa.gov/air_traffic/publications/atpubs/fs_html/chap11_section_1.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["FAA operations prescribe standardized words and phrases for radiotelephone and interphone communication.","The ICAO phonetic alphabet and specified number pronunciations are used for clarity.","FAA explicitly specifies procedures for relaying ATC communications, demonstrating that codebook and relay wording are governed operational artifacts."]},{"source_id":"S7","title":"Prediction of word confusabilities for speech recognition","publisher":"International Speech Communication Association","url":"https://www.isca-archive.org/icslp_1994/roe94_icslp.html","source_class":"PRIMARY_RESEARCH","publication_date":"1994","accessed_at":"2026-08-03","claims_supported":["Prior research built a tool that detects confusable vocabulary using pronunciation and empirical phonetic-confusion information.","The proposed use was to detect and eliminate confusable words from speech-recognition vocabularies.","This establishes pairwise confusability-based vocabulary redesign as prior art."]},{"source_id":"S8","title":"Minimizing Sequential Confusion Error in Speech Command Recognition","publisher":"International Speech Communication Association","url":"https://www.isca-archive.org/interspeech_2022/yang22m_interspeech.html","source_class":"PRIMARY_RESEARCH","publication_date":"2022","accessed_at":"2026-08-03","claims_supported":["Command confusion among similar pronunciations is a recognized speech-command problem.","The authors optimized a sequential-confusion loss using constructed confusing-command sets.","Their experiment reported an 18.28% reduction in confusion errors and a 33.7% relative false-reject-rate reduction at the stated operating point.","Direct recognizer training against command confusion is an existing comparator distinct from codebook spectral redesign."]}],"problem_evidence":{"support":"STRONG","rationale":"Operational FAA data show readback errors and misunderstandings; EUROCONTROL treats communication confusion as a hazard; controlled transmission-chain experiments show verbal information degrades and that some content persists more strongly than other content. NASA further documents noise, speech variability, task loading, command substitution, and operational degradation in voice-command systems. These sources support the general problem and the possibility of structured relay persistence, but none demonstrates the candidate's specific invariant eigenmode in a fixed safety-command alphabet.","source_ids":["S1","S2","S3","S4"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"EUROCONTROL explicitly called for implementation of communication-safety recommendations, while NASA supplies a concrete adopter class—spaceflight command-and-control project teams—and requirements for vocabulary testing, confirmation, standardized nomenclature, and safe fallback. FAA likewise governs phraseology and relayed communications. No source identifies an organization requesting eigendecomposition of a command-confusion operator or commits data, funding, or deployment authority to this candidate.","source_ids":["S2","S4","S5","S6"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"NASA voice-controlled spacecraft development workflow","similarity":"Uses a complete vocabulary confusability matrix, vocabulary replacement, recognition-error handling, confirmation dialogue, operational-condition testing, task-level command accuracy, and human-centered approval considerations.","remaining_difference":"It evaluates cells, vocabulary items, and dialogue states rather than eigenmodes of a repeated-relay operator, and it does not select a minimal redesign by damping action-weighted persistent modes.","source_ids":["S4"]},{"name":"Pairwise word-confusability vocabulary screening","similarity":"Predicts which command words will be confused and supports eliminating or replacing confusable words before deployment.","remaining_difference":"It is pairwise and pronunciation-oriented; it neither models repeated relay nor ranks coupled, alphabet-wide invariant directions.","source_ids":["S7"]},{"name":"Minimizing Sequential Confusion Error training","similarity":"Directly optimizes a speech-command recognizer against confusing command sequences and empirically reduces command-confusion errors.","remaining_difference":"It modifies model training rather than the governed codebook and confirmations, and its sequence criterion is not spectral persistence under repeated human-machine relay.","source_ids":["S8"]},{"name":"Standard phraseology and ICAO phonetic codewords","similarity":"Redesigns spoken signal forms and prescribes relay wording to reduce ambiguity in noisy operational radio communication.","remaining_difference":"The codewords and phraseology are standardized globally rather than selected from an estimated local confusion operator or action-specific modal loss.","source_ids":["S6"]},{"name":"Universal confirmation or repetition for critical commands","similarity":"Adds a closed-loop check to prevent hazardous command errors; experiments also show repetition can improve spoken relay accuracy.","remaining_difference":"It confirms based on command criticality or uniformly, rather than targeting checks to commands loading on a persistent hazardous mode under a matched latency budget.","source_ids":["S3","S5"]}],"distinctive_claim_remaining":"For one frozen, approximately stationary command-relay path, selecting at most a predeclared small number of token changes and targeted confirmation prompts by the conditioned nontrivial eigenmodes of its empirical confusion operator will reduce held-out, action-weighted multi-relay substitution loss more than unchanged vocabulary, worst-command renaming, pairwise-confusability design, universal confirmation, and direct expected-cost optimization under identical vocabulary-change and latency budgets. The claim is falsified if the single-operator model fails, modes are ill-conditioned or unstable across speakers and conditions, or a comparator matches or exceeds every declared safety and latency outcome.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Confusion-matrix collection, controlled vocabulary testing, codeword replacement, confirmation dialogue, and held-out evaluation are established and technically implementable. Matrix eigendecomposition itself is routine. The central unverified assumptions are that a heterogeneous radio-ASR-human relay can be represented by one stationary square operator, that its nontrivial eigenvectors are sufficiently conditioned and stable to interpret, and that damping their gains improves action-weighted outcomes beyond direct optimization. Privacy/consent, accent and speech-variation coverage, configuration control, training, fallback, and safety approval are manageable but partner-specific. No reviewed source validates the complete spectral workflow in this use case.","source_ids":["S3","S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Command substitutions in aviation, emergency relay, and spacecraft control can be hazardous, and official sources treat communication accuracy and confirmation as safety-relevant. Realized impact for this method is unmeasured.","source_ids":["S1","S2","S4","S5"]},"stakeholder_pull":{"score":3,"rationale":"NASA, EUROCONTROL, and FAA visibly govern and invest in safer voice communications, but no named stakeholder has requested or sponsored the spectral method.","source_ids":["S2","S4","S5","S6"]},"incremental_advantage":{"score":3,"rationale":"Coupled persistent-mode targeting could outperform cell-wise or pairwise redesign when relay effects accumulate, but the advantage requires controlled head-to-head evidence and may disappear against direct cost optimization or confirmation.","source_ids":["S3","S4","S7","S8"]},"distinctiveness_plausibility":{"score":3,"rationale":"The searched prior art covers confusion matrices, vocabulary replacement, sequence-confusion training, standardized codewords, and confirmation, but no close source applied conditioned nontrivial eigenmodes to minimal relayed-command codebook redesign. This is a bounded-search distinction, not a world-novelty finding.","source_ids":["S4","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"The data collection and computational components are conventional and can be run offline. Heterogeneous-step operators, rare-event sparsity, non-normality, and mode instability are substantial but explicitly testable limitations.","source_ids":["S3","S4","S7","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"NASA and FAA examples show that vocabulary, confirmation, and phraseology have identifiable program or regulatory owners. The candidate lacks a named operational sponsor, local safety authority, and approved change-control path.","source_ids":["S4","S5","S6"]},"evidence_readiness":{"score":3,"rationale":"A controlled corpus and clear comparators can be assembled, but the decisive evidence needs new speaker-stratified relay trials rather than public web data.","source_ids":["S1","S3","S4"]},"safety_net_benefit":{"score":4,"rationale":"The method can remain advisory while existing vocabulary, mandatory confirmations, alternative controls, and human override remain intact; these controls reduce first-test risk.","source_ids":["S4","S5"]},"scalability":{"score":3,"rationale":"The analytical software can transfer across bounded vocabularies, but every materially different channel, speaker population, decoder version, and relay procedure requires re-estimation and approval.","source_ids":["S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One offline study covering one 8–20-command vocabulary, approximately 24 consenting speakers, four noise/channel strata, relay depths one through four, frozen transcription software, analysis, preregistration, and a shadow-trial report.","confidence":"LOW","assumptions":["Existing recording and usability facilities are available.","No live operational traffic or certified-system modification is required.","The estimate is resource-equivalent labor and facility use, not a vendor quote.","Rare hazardous commands do not require an orders-of-magnitude larger sample."],"source_ids":["S3","S4"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Partner-specific protocol engineering, larger representative validation, transcription integration, configuration management, safety and human-factors review, training materials, and backward-compatible codebook/version design.","confidence":"LOW","assumptions":["A single service and one primary language are in scope.","Existing radios and transcription infrastructure remain in place.","No formal aircraft or spacecraft recertification campaign is included.","Mandatory confirmation and fallback controls are retained."],"source_ids":["S4","S5","S6"]},"operational_launch":{"band_2026_usd":"1M_TO_5M","scope":"Authorized rollout for one safety-critical organization, including multisite validation, operator training, parallel-version controls, cutover exercises, documentation, assurance review, and rollback readiness.","confidence":"LOW","assumptions":["Launch spans multiple teams or sites but not an international standard change.","Mixed-codebook operation requires explicit transition controls.","The band excludes replacement of radio hardware or the core transcription platform.","Formal authority effort could move the cost outside this band."],"source_ids":["S2","S4","S5","S6"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Periodic corpus refresh, drift and residual testing after channel or decoder changes, safety review, version governance, refresher training, and incident analysis for one organization.","confidence":"LOW","assumptions":["Two to four formal reevaluations or change-triggered reviews occur annually.","A standing communications, data-science, human-factors, and safety team shares the work.","No major platform replacement occurs."],"source_ids":["S4","S5"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official operational evidence and primary relay experiments establish consequential voice-communication errors, structured persistence, and safety relevance, although not the proposed eigenmode mechanism itself.","source_ids":["S1","S2","S3","S4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"NASA command-and-control programs, FAA phraseology authorities, and EUROCONTROL safety stakeholders are identifiable adopter or authorizer classes with expressed need for vocabulary testing, confirmation, standardization, and communication-safety improvement. No specific organization has committed to this proposal.","source_ids":["S2","S4","S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The candidate can be contrasted against established confusion-matrix testing, pairwise vocabulary screening, direct sequence-confusion optimization, standard codewords, and universal confirmation under matched budgets.","source_ids":["S4","S5","S6","S7","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"One frozen vocabulary, channel, decoder, participant sample, relay-depth range, redesign budget, and held-out comparison can test both the modal premise and intervention advantage without operational deployment.","source_ids":["S3","S4"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step can be limited to consenting participants, simulated impairments, frozen data, and non-operational shadow analysis while mandatory confirmation, alternative control, human override, and the existing codebook remain unchanged. This does not authorize deployment.","source_ids":["S4","S5"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Resource-equivalent bands can be scoped, but the candidate lacks a named organization, vocabulary size, assurance regime, sample-size calculation, vendor integration plan, and certification pathway; no source provides a directly applicable cost basis.","source_ids":["S4","S5","S6"]}},"next_evidence_step":"Pre-register a non-operational study of one frozen 8–20-command alphabet with approximately 24 consenting speakers, four declared noise/channel strata, one frozen transcription version, and relay depths one through four. Split by speaker and condition before fitting. Estimate the one-step operator only on the fitting partition; bootstrap eigenvalues, invariant-subspace angles, condition numbers, and spectral gaps; and compare a homogeneous single-operator model with depth-specific and ordered-product alternatives. Permit at most three token changes and one targeted confirmation selected without held-out outcomes. On held-out chains, compare unchanged vocabulary, universal confirmation, worst-command renaming, pairwise-confusability redesign, and direct action-cost optimization under identical change-count and latency budgets. The primary endpoint is action-weighted hazardous substitution loss at the maximum relay depth; secondary endpoints are minimum per-command accuracy, latency, residual structure, subgroup performance, modal-gain change, and basis stability. Falsify the premise if no reproducible nontrivial mode exists, if operator homogeneity or conditioning thresholds fail, or if multi-relay errors are adequately predicted by coordinate-level measures. Falsify the intervention if any budget-matched comparator equals or improves every declared safety endpoint, or if the redesign worsens a hazardous substitution, subgroup floor, latency limit, or residual criterion. A positive result may authorize only a larger partnered shadow trial.","blocking_evidence":["No direct source was found validating eigenmode-guided codebook or confirmation redesign for a relayed finite spoken-command alphabet.","A named operational partner, data owner, safety authorizer, and fixed command vocabulary have not been identified.","Stationarity, Markov sufficiency, and use of one common operator across radio, transcription, and human-relay steps are untested.","The frequency of rare hazardous substitutions and resulting sample-size requirement are unknown.","Mode conditioning, bootstrap stability, and invariance across speakers, accents, noise, relay depth, and transcription versions are unknown.","No partner-specific privacy, labor, accessibility, configuration-control, certification, or regulatory assessment has been completed.","Cost bands are planning estimates without vendor quotes, staffing rates, certification scope, or an operational rollout footprint.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes adjacent practices and a bounded contrast only. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named communications-system partner, safety authorizer, fixed vocabulary, approved corpus protocol, and data-governance plan.","Demonstrate out-of-sample that a stationary or explicitly ordered relay operator predicts depth-two through depth-four substitutions better than independent one-hop command metrics.","Predeclare and pass eigenvalue-conditioning, invariant-subspace stability, spectral-gap, residual, subgroup, and minimum-command-accuracy thresholds.","Show lower held-out action-weighted substitution loss than all matched-budget comparators without increasing hazardous substitutions or violating latency and confirmation requirements.","Produce partner-specific startup, launch, recurring-cost, training, version-transition, fallback, and authority estimates."],"reason":"Web evidence verifies the broader safety problem, credible authorizer classes, adjacent prior art, and a falsifiable contrast, but it cannot establish stable relay modes or comparative intervention benefit. Those questions require new controlled human/ASR relay data and live partner authority. Under the controller rule, that evidence dependency requires STOP_EMPIRICAL_RESEARCH_NEEDED; repairable is false as required for every STOP recommendation."},"proposal_index":4}