{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__veterinary_medicine:P5:v0","cell_id":"computability_boundary_mapping__veterinary_medicine","search_queries":["site:avma.org veterinary behavior assessment diagnosis history observation official","site:acvb.org behavior assessment veterinarian diagnosis animal behavior official","adaptive distinguishing sequence finite state machines synthesis paper","active automata learning adaptive distinguishing sequences LearnLib","veterinary behavior diagnosis assessment history observation site:edu veterinary behavior service","site:avma.org animal welfare principles veterinary procedures pain distress official","site:aphis.usda.gov Animal Welfare Act IACUC review procedures animals official","veterinary behavior assessment reliability diagnosis study animal behavior","animal behavior adaptive experiment design model discrimination study","veterinary behavior optimal experimental design discriminate models animal response","active learning animal behavior experiment adaptive testing models","model discrimination optimal experimental design biological systems review","\"What is the evidence for reliability and validity of behavior evaluations for shelter dogs\"","\"Effective algorithms for constructing minimum cost adaptive distinguishing sequences\" DOI","site:pubmed.ncbi.nlm.nih.gov shelter dog behavior evaluations reliability validity Patronek","adaptive distinguishing sequences finite state machines Lee Yannakakis PDF"],"sources":[{"source_id":"S1","title":"Behavior Medicine","publisher":"Cornell University College of Veterinary Medicine","url":"https://www.vet.cornell.edu/hospitals/services/behavior","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Cornell operates an identifiable veterinary Behavior Medicine Service.","Its workflow reviews medical and behavioral history, performs risk assessment, and discusses diagnosis, prognosis, and treatment.","Physical examination is limited by unnecessary stress, and behavior videos are requested only when safe."]},{"source_id":"S2","title":"Can This Dog Be Rehomed to You? A Qualitative Analysis and Assessment of the Scientific Quality of the Potential Adopter Screening Policies and Procedures of Rehoming Organisations","publisher":"Frontiers in Veterinary Science","url":"https://www.frontiersin.org/journals/veterinary-science/articles/10.3389/fvets.2020.617525/pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2020-12","accessed_at":"2026-08-03","claims_supported":["Survey data from 82 rehoming organizations showed substantial variation in assessment processes and limited quality-control evidence.","Only eight of 31 potentially exclusionary adopter characteristics had potential support in the reviewed evidence.","The study documents broader uncertainty and weak scientific support around consequential animal-related assessments, but not demand for executable-model protocol synthesis."]},{"source_id":"S3","title":"Animal Welfare Act and Animal Welfare Regulations","publisher":"USDA Animal and Plant Health Inspection Service","url":"https://www.aphis.usda.gov/media/document/17164/file","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2023-05-01","accessed_at":"2026-08-03","claims_supported":["For covered research activities, an IACUC may approve, require modification of, withhold approval of, or suspend animal activities.","Covered procedures must avoid or minimize discomfort, distress, and pain, and relevant planning may require attending-veterinarian consultation.","Institutional officials cannot approve covered animal activity lacking IACUC approval; qualifying field studies are exempt from the cited IACUC requirement."]},{"source_id":"S4","title":"Active Automata Learning with Adaptive Distinguishing Sequences","publisher":"arXiv / TU Dortmund University","url":"https://arxiv.org/abs/1902.01139","source_class":"PRIMARY_RESEARCH","publication_date":"2019-02-04","accessed_at":"2026-08-03","claims_supported":["Adaptive distinguishing trees were already developed for active automata learning.","The ADT algorithm was integrated into the open-source LearnLib framework.","Adaptive action choices based on prior outputs are established formal-methods prior art."]},{"source_id":"S5","title":"Incomplete Adaptive Distinguishing Sequences for Non-deterministic FSMs","publisher":"IEEE Transactions on Software Engineering / White Rose Research Online","url":"https://eprints.whiterose.ac.uk/id/eprint/200869/","source_class":"PRIMARY_RESEARCH","publication_date":"2023-07-05","accessed_at":"2026-08-03","claims_supported":["Prior research treats adaptive distinguishing tests for nondeterministic, partial, observable finite-state machines.","Not every system possesses an adaptive distinguishing sequence.","Checking for sets of incomplete adaptive distinguishing sequences is PSPACE-hard, and generation algorithms have been implemented and experimentally evaluated."]},{"source_id":"S6","title":"LearnLib: An Open Framework for Automata Learning","publisher":"LearnLib Project, TU Dortmund University","url":"https://learnlib.de/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2025-02-06","accessed_at":"2026-08-03","claims_supported":["A maintained open-source product already provides active automata-learning algorithms, including ADT.","The product includes complete depth-bounded exploration, random and W-method-style equivalence approximations, caches, parallelism, and model abstractions.","Existing software substantially lowers implementation barriers for bounded automata experiments."]},{"source_id":"S7","title":"State Identification for Labeled Transition Systems with Inputs and Outputs","publisher":"arXiv / Radboud University","url":"https://arxiv.org/abs/1907.11034","source_class":"PRIMARY_RESEARCH","publication_date":"2019-07-25","accessed_at":"2026-08-03","claims_supported":["Algorithms already synthesize adaptive tests for richer input/output transition systems with observable nondeterminism and partiality.","The work formalizes adversarial compatibility and proves that its algorithm finds a distinguishing test when one exists under its assumptions.","The paper reports exponential worst-case structure and explicitly separates complete discrimination from high but incomplete empirical coverage."]},{"source_id":"S8","title":"Adaptive Optimal Training of Animal Behavior","publisher":"Neural Information Processing Systems Foundation","url":"https://papers.nips.cc/paper_files/paper/2016/hash/7fec306d1e665bc9c748b5d2b99a6e97-Abstract.html","source_class":"PRIMARY_RESEARCH","publication_date":"2016","accessed_at":"2026-08-03","claims_supported":["Adaptive computational selection of stimuli has already been applied to animal behavior.","The study used rat sensory-discrimination data to infer learning-model parameters and select subsequent training stimuli.","This is animal-domain adjacent prior art, although it optimizes learning rather than guaranteeing worst-case separation of two finite models."]}],"problem_evidence":{"support":"WEAK","rationale":"Veterinary and rehoming organizations visibly conduct consequential behavioral and risk assessments, and published research documents weak validation and quality-control problems. No source found evidence that veterinary or wildlife-rehabilitation teams are pursuing a universal executable-model planner, interpreting synthesis timeouts as biological indistinguishability, or need a computability-boundary intervention. The proposal therefore maps a real assessment-validity concern onto an externally unverified computational failure mode.","source_ids":["S1","S2"]},"stakeholder_evidence":{"support":"WEAK","rationale":"Cornell's Behavior Medicine Service is an identifiable potential adopter with a real assessment workflow. For covered research-animal use, an IACUC and attending veterinarian are identifiable authorizers. Neither expresses demand for bounded adaptive-protocol synthesis, formal response models, exhaustive certificates, or computability analysis; field-study and clinical authority also differ from the cited federal research rules.","source_ids":["S1","S3"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Adaptive distinguishing trees in LearnLib","similarity":"Already constructs adaptive input/output discrimination structures and provides depth-bounded exploration in maintained software.","remaining_difference":"It is framed for automata learning and conformance testing, not veterinary model-pair assessment, welfare routing, or the proposal's four-result vocabulary.","source_ids":["S4","S6"]},{"name":"State identification for input/output labeled transition systems","similarity":"Synthesizes adaptive tests under output nondeterminism, adversarial compatibility, partial inputs, and formal completeness conditions.","remaining_difference":"It distinguishes states in a transition system rather than two separately packaged animal-response hypotheses and does not include veterinary governance.","source_ids":["S7"]},{"name":"Incomplete adaptive distinguishing sequences for nondeterministic FSMs","similarity":"Directly addresses existence and construction of adaptive distinguishing tests for finite nondeterministic and partial observable machines.","remaining_difference":"Its target is software-model state identification and incomplete sequence sets, not a welfare-reviewed assessment between biological hypotheses.","source_ids":["S5"]},{"name":"Adaptive optimal training of animal behavior","similarity":"Uses learned animal-response models to adaptively choose stimuli in a real animal experiment.","remaining_difference":"It optimizes expected training progress and parameter inference rather than providing exhaustive worst-case discrimination or bounded negative certificates.","source_ids":["S8"]}],"distinctive_claim_remaining":"A contrastive claim remains only at the integration layer: for two explicitly finite veterinary response models, a welfare-gated wrapper around established adaptive-testing methods can preserve every declared state/outcome path, emit replayable positive or exhaustive bound-qualified negative evidence, and prevent UNKNOWN or OUT-OF-MODEL from being communicated as biological indistinguishability. This is falsifiable but is not presently supported as a novel algorithmic contribution.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Finite adaptive-test synthesis, nondeterministic transition-system handling, and depth-bounded exploration have algorithms and working software. It is a reasonable inference that two model hypotheses can be encoded in a tagged combined transition system for a small synthetic prototype. Evidence does not establish biological model fidelity, usability of the certificates, tractability at useful veterinary bounds, or live-animal effectiveness.","source_ids":["S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Bound-qualified results could reduce overstatement in consequential behavior assessments, but the specific timeout-to-indistinguishability failure is not externally documented.","source_ids":["S1","S2"]},"stakeholder_pull":{"score":2,"rationale":"A plausible veterinary adopter and research authorizer exist, but neither requests this intervention or formal-model workflow.","source_ids":["S1","S3"]},"incremental_advantage":{"score":2,"rationale":"Certificates and explicit result routing improve governance over heuristic search, but the computational capability substantially overlaps established adaptive-testing methods and tools.","source_ids":["S4","S5","S6","S7"]},"distinctiveness_plausibility":{"score":1,"rationale":"The bounded synthesis core, adaptive decision structure, nondeterministic semantics, and completeness qualification are established; veterinary terminology and governance are the principal remaining differences.","source_ids":["S4","S5","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"The proposed tiny synthetic cases are readily implementable using known finite-state algorithms or exhaustive enumeration, although scaling and certificate completeness require careful verification.","source_ids":["S5","S6","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"Simulation-only work needs little special authority; covered live research has a defined IACUC and veterinary pathway. Clinical and wildlife contexts require separately resolved authority.","source_ids":["S1","S3"]},"evidence_readiness":{"score":4,"rationale":"The implementation and label contract can be tested immediately on small synthetic models with independent brute-force comparators and no animals.","source_ids":["S5","S6","S7"]},"safety_net_benefit":{"score":3,"rationale":"Explicit UNKNOWN, OUT-OF-MODEL, and bound-qualified negatives are sensible safeguards, but no evidence quantifies current mislabeling frequency or downstream harm reduction.","source_ids":["S1","S2","S3"]},"scalability":{"score":2,"rationale":"Decision-tree and product-state growth can be exponential, and useful biological formalizations would be bespoke; existing tools mitigate engineering effort but not state explosion.","source_ids":["S5","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"UNDER_10K","scope":"Implement the two tiny synthetic model pairs, exhaustive tree enumerator, independent path checker, replayable certificates, OUT-OF-MODEL case, and forced UNKNOWN case; conduct one independent code/proof review.","confidence":"MODERATE","assumptions":["One experienced formal-methods engineer for approximately one to two weeks.","One reviewer for one to two days.","No live animals, clinical data, paid software, regulatory submission, or production integration.","Existing open-source automata components may be reused."],"source_ids":["S6","S7"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production-quality schema validation, certificate verifier, audit records, tests, security review, user interface, and co-design with one veterinary behavior service.","confidence":"LOW","assumptions":["Two to four technical and domain contributors over approximately three to six months.","Existing clinic infrastructure can accept a standalone prototype.","Biological model creation remains limited to a small pilot domain.","No live-animal protocol execution is included."],"source_ids":["S1","S6"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Launch across multiple veterinary or research teams with model-authoring support, governance, documentation, training, validation, and any necessary animal-use review.","confidence":"LOW","assumptions":["Substantial domain-expert effort is needed to formalize observations, stimuli, state transitions, and welfare exclusions.","Covered live research requires institutional review and veterinary participation.","The launch does not include major sensors, laboratory construction, or a clinical trial."],"source_ids":["S1","S3"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Software maintenance, model/version review, certificate audits, user support, governance refresh, and periodic revalidation of approved use cases.","confidence":"LOW","assumptions":["A small user base and limited number of model families.","Approximately 0.5 to 1.5 technical/domain full-time-equivalent resources.","No high-volume live-animal study operations are included."],"source_ids":["S1","S3","S6"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"External evidence supports consequential behavioral-assessment uncertainty, but not the proposal's specific universal-planner, timeout, or computability-boundary failure mode.","source_ids":["S1","S2"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Cornell is a credible veterinary behavior-service adopter candidate, and IACUCs plus attending veterinarians are credible authorizers for covered research-animal activities, although expressed demand for this tool is absent.","source_ids":["S1","S3"]},"distinct_testable_incremental_claim":{"status":"NO","reason":"The testable bounded guarantee substantially collides with established adaptive distinguishing-sequence and input/output transition-system practice; the remaining veterinary wrapper is not shown to provide a distinct technical advantage.","source_ids":["S4","S5","S6","S7","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A simulation-only experiment with small finite models, exhaustive comparators, known positive and negative cases, and explicit routing falsifiers is tightly bounded.","source_ids":["S5","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The authorized first step is synthetic and uses no animals or clinical records. Any later covered animal activity has an explicit review pathway and must remain outside the prototype authorization.","source_ids":["S3"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands are broad resource-equivalent estimates tied to explicit labor, reuse, review, integration, and animal-use assumptions; confidence remains low beyond the synthetic prototype because no vendor quotes or adopter-specific workflow data were found.","source_ids":["S1","S3","S6"]}},"next_evidence_step":"Run a simulation-only differential benchmark on two pairs of finite response models, each limited to four states per model, three actions, three observations, two outcomes per transition, and depth three. Compare (A) the proposed exhaustive paired-model enumerator, (B) an established adaptive-distinguishing/splitting-graph implementation or faithful reimplementation, and (C) independent brute-force enumeration. Require an agreed exact count of admissible trees and paths, replayable positive certificates, a complete bound-qualified negative, rejection of one executable out-of-contract model, and UNKNOWN under forced interruption. Falsify the implementation if any admissible path or tree is omitted, any certificate fails replay, algorithms disagree without resolution, or any interrupted/out-of-model result is reported as NONE. This step tests implementation and labeling only, not biological validity, stakeholder demand, live-animal welfare, or practical impact.","blocking_evidence":["No externally documented veterinary or wildlife-rehabilitation demand for universal executable-model discrimination or bounded computability mapping.","No independently checked reduction establishes undecidability for the proposal's exact unrestricted encoding and quantifiers.","No evidence shows that finite animal-response models preserve clinically or behaviorally important reality.","No field evidence shows that certificates and four-way result labels improve veterinary decisions or prevent actual timeout misinterpretation.","The core bounded synthesis capability substantially overlaps established adaptive distinguishing-test algorithms and maintained tooling.","Deployment and recurring cost estimates lack adopter-specific labor measurements, procurement quotes, and integration requirements."],"research_disposition":"KNOWN_PRACTICE_DIFFUSION","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact were not measured. The search establishes substantial collision with accessible adaptive distinguishing-sequence, nondeterministic transition-system, automata-learning, and adaptive animal-experiment literature; it does not establish that no veterinary-specific implementation, unpublished system, patent, or exact paired-model formulation exists elsewhere.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_DIFFUSION","repairable":false,"material_progress_observed":false,"progress_targets":["Reclassify the concept as veterinary adaptation and governed diffusion of established adaptive-testing methods, not as a distinct bounded-synthesis innovation.","Obtain a written problem statement from at least one named veterinary behavior or wildlife-rehabilitation service documenting actual model-discrimination decisions, timeout errors, desired actions and observations, and decision authority.","Benchmark the proposed wrapper directly against established ADS, splitting-graph, and LearnLib capabilities and retain only differences that survive equivalent encodings and certificate checks.","Either produce an independently checked reduction for the exact unrestricted target or remove the undecidability claim from public language.","Before any live use, establish biological model validity, welfare constraints, and the applicable clinical, institutional, and regulatory authorization pathway."],"reason":"The broader assessment-validity concern is real and the synthetic implementation is feasible, but bounded adaptive discrimination for finite nondeterministic input/output models is established practice with available tooling. The remaining contribution is a veterinary workflow and governance translation for which expressed demand and biological validity are unverified. This is therefore a diffusion opportunity, not a distinct innovation candidate."},"proposal_index":5}