{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"computability_boundary_mapping__engineering_design:RETRIEVAL_FIRST:v0","cell_id":"computability_boundary_mapping__engineering_design","search_queries":["site:snl-dakota.github.io simulation failure capturing Dakota retry recovery continuation","site:fmi-standard.org docs status discard error output undefined FMI 3.0","site:openmdao.org AnalysisError optimizer incorrect solution","site:nasa.gov systems engineering handbook decision analysis assumptions uncertainty rationale","surrogate optimization computationally expensive black-box hidden constraints failed evaluations primary research PDF","human in the infinite loop inconsistent contradictory human judgments Bayesian optimization PDF","site:mathworks.com optimizing simulation objective function NaN failed evaluation solver","BLS software developers median pay May 2025 engineers hourly compensation"],"sources":[{"source_id":"S1","title":"Simulation Failure Capturing — Dakota 6.19.0 Documentation","publisher":"Sandia National Laboratories","url":"https://snl-dakota.github.io/docs/6.19.0/users/usingdakota/advanced/simulationfailurecapturing.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"n.d.; version 6.19.0","accessed_at":"2026-08-02","claims_supported":["External simulation failures are a recognized optimization-workflow problem requiring detection, communication, and mitigation.","Dakota implements abort, retry, dummy-value recovery, and continuation routes.","Retry covers unavailable licenses, networking, and shared-storage failures.","Dummy objective or constraint values can intentionally convert failed evaluations into optimizer-visible numerical penalties."]},{"source_id":"S2","title":"Functional Mock-up Interface Specification 3.0.1","publisher":"Modelica Association Project FMI","url":"https://fmi-standard.org/docs/3.0.1/","source_class":"STANDARD","publication_date":"2022","accessed_at":"2026-08-02","claims_supported":["FMI defines typed OK, Warning, Discard, Error, and Fatal outcomes for external model calls.","Discard and Error make output arguments undefined rather than valid numerical evidence.","Discard permits only controlled alternatives or termination; Error prohibits continuing the simulation without restoration or reset.","Unavailable required runtime resources must be logged and returned as an error."]},{"source_id":"S3","title":"Raising an AnalysisError","publisher":"OpenMDAO Development Team","url":"https://openmdao.org/newdocs/versions/latest/advanced_user_guide/analysis_errors/analysis_error.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026 online edition","accessed_at":"2026-08-02","claims_supported":["OpenMDAO exposes invalid analysis regions through a distinct AnalysisError.","Optimizer reactions differ, and some optimizers can fail or return an incorrect solution after analysis errors.","Correct handling therefore depends on the optimizer and workflow, not merely on emitting an exception."]},{"source_id":"S4","title":"Optimizing a Simulation or Ordinary Differential Equation","publisher":"MathWorks","url":"https://www.mathworks.com/help/optim/ug/optimizing-a-simulation-or-ordinary-differential-equation.html","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-02","claims_supported":["MathWorks recommends returning NaN when an objective or constraint evaluation fails so supported solvers can try another step.","MathWorks warns that returning an arbitrary large value can confuse a solver because the value appears genuine.","Failed-evaluation propagation is established commercial practice, although it does not preserve an unresolved state through an engineering gate."]},{"source_id":"S5","title":"Surrogate Optimization of Computationally Expensive Black-Box Problems with Hidden Constraints","publisher":"INFORMS Journal on Computing","url":"https://pubsonline.informs.org/doi/10.1287/ijoc.2018.0864","source_class":"PRIMARY_RESEARCH","publication_date":"2019","accessed_at":"2026-08-02","claims_supported":["Hidden constraints arise when black-box objective evaluation returns no value for a parameter vector.","The SHEBO method predicts evaluability and adjusts an evaluability threshold, showing that missing simulator outputs are an established optimization research problem.","The paper addresses simulation evaluability, not end-to-end heterogeneous evaluator contracts at design gates."]},{"source_id":"S6","title":"The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures","publisher":"Association for Computing Machinery authors via arXiv","url":"https://arxiv.org/abs/2207.12761","source_class":"PRIMARY_RESEARCH","publication_date":"2022-07-26","accessed_at":"2026-08-02","claims_supported":["A three-month professional deployment found inconsistent and contradictory human judgments in preference-guided Bayesian optimization.","The studied optimization lacked mechanisms for handling those contradictions.","Human ratings violated stable-utility assumptions and propagated errors into system outcomes, supporting the problem beyond numerical simulators."]},{"source_id":"S7","title":"NASA Systems Engineering Handbook, Section 6.8: Decision Analysis","publisher":"National Aeronautics and Space Administration","url":"https://www.nasa.gov/reference/6-8-decision-analysis/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2023 online edition","accessed_at":"2026-08-02","claims_supported":["NASA identifies a decision authority or decision-making body as the recipient of engineering analysis and recommendations.","NASA says tool assumptions and limitations must be understood, documented, and integrated with other decision factors.","NASA calls for sufficient technical evidence and characterization of uncertainty and for reports containing criteria, methods, results, recommendations, and final decisions.","This identifies NASA program decision authorities as a credible authorizer class with an expressed need for qualified, traceable decision evidence."]},{"source_id":"S8","title":"Occupational Employment and Wages News Release — May 2025 Estimates","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-02","claims_supported":["May 2025 mean wages were $71.20 per hour for software developers and $52.84 per hour for industrial engineers.","The wage estimates provide an official labor-cost anchor for broad 2026 resource-equivalent bands.","The published figures are wages rather than fully loaded project rates, so overhead, benefits, management, and vendor costs remain assumptions."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible across independent settings. Dakota documents external simulation and infrastructure failures; OpenMDAO warns that analysis errors can lead to failure or an incorrect solution; MathWorks warns that fabricated large values can confuse solvers; hidden-constraint research treats absent simulation outputs as a recurring optimization problem; and a professional human-in-the-loop study observed inconsistent judgments that propagated into outcomes. The sources do not quantify how often multidisciplinary design gates issue false-definitive verdicts, so prevalence and realized harm remain unmeasured.","source_ids":["S1","S3","S4","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"NASA guidance identifies the decision authority and explicitly requires sufficient technical evidence, uncertainty characterization, and documentation of analytical assumptions and limitations. Dakota, FMI, OpenMDAO, and MathWorks also demonstrate practitioner demand for failed-evaluation semantics. No named organization was found requesting or committing to adopt the proposed four-evaluator, end-to-end router, so this is credible authorizer need rather than demonstrated procurement pull.","source_ids":["S1","S2","S3","S4","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Dakota simulation-failure capturing","similarity":"Implements explicit failure communication and policy-selected abort, retry, dummy-value recovery, or continuation for external analyses, closely matching the router's executable core.","remaining_difference":"It is simulation-centered, permits dummy numerical recovery, and does not impose a heterogeneous evaluator contract or preserve UNRESOLVED through a multidisciplinary design gate.","source_ids":["S1"]},{"name":"FMI typed status and undefined-output contract","similarity":"Provides standardized typed success and failure states, controlled recovery rules, and undefined-output semantics for failed external model calls.","remaining_difference":"It governs model interoperability rather than optimizer verdicts, laboratory or human evaluators, mechanically checked promises, or board escalation.","source_ids":["S2"]},{"name":"OpenMDAO AnalysisError and MathWorks NaN handling","similarity":"Both expose unevaluable analyses to solvers and allow controlled movement away from failed points instead of treating failures as ordinary evidence.","remaining_difference":"Neither examined source carries a typed unresolved state across four evaluator classes into a final design-gate invariant.","source_ids":["S3","S4"]},{"name":"SHEBO hidden-constraint optimization","similarity":"Treats missing black-box simulation values as evaluability failures and uses an explicit evaluability model.","remaining_difference":"It predicts simulator evaluability rather than enforcing promises and authority-preserving escalation for heterogeneous evidence providers.","source_ids":["S5"]},{"name":"NASA decision-analysis reporting practice","similarity":"Requires decision criteria, methods, evidence, uncertainty, assumptions, limitations, recommendations, and final-decision provenance for a decision authority.","remaining_difference":"It is governance guidance rather than an executable contract that blocks unqualified optimization labels when required evidence is unresolved.","source_ids":["S7"]}],"distinctive_claim_remaining":"On a prespecified multidisciplinary trace set, a single mechanically enforced contract spanning simulator, laboratory, vendor, and human evaluators can reduce the rate at which required non-results become unqualified feasible, infeasible, or optimal gate verdicts to zero, relative to the existing workflow and a dummy-value/NaN-style comparator, without blocking any valid in-promise success. No examined source demonstrated that exact four-class, end-to-end integration; the performance claim remains untested.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Typed adapter states, retry/abort routing, undefined-output semantics, optimizer error propagation, audit records, and decision-authority workflows are individually implemented or officially specified, making a one-workflow shadow prototype technically plausible. Unverified portions are mechanical promise checks for laboratories, vendors, and humans; propagation through every downstream interface; acceptable escalation load; access to proprietary traces; vendor-contract, privacy, records-management, and security constraints; and production authority. The design-review board must retain decision authority, and the first test must remain non-decisional and reversible.","source_ids":["S1","S2","S3","S4","S7"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Preventing unsupported feasibility or optimality conclusions can protect consequential engineering decisions, but prevalence, loss magnitude, and realized impact have not been measured.","source_ids":["S3","S4","S6","S7"]},"stakeholder_pull":{"score":3,"rationale":"Official guidance expresses strong need for qualified evidence and explicit uncertainty, while product documentation confirms practitioner demand for failure handling; no adopter commitment or budget was found.","source_ids":["S1","S2","S7"]},"incremental_advantage":{"score":3,"rationale":"End-to-end preservation of unresolved status across heterogeneous evaluator types would improve guarantee traceability over solver-local failure handling, but most constituent mechanisms already exist.","source_ids":["S1","S2","S3","S4","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"Prior-art collision is substantial. Only the narrow integration across four evaluator classes and the final-gate invariant remains plausibly distinctive, and ordinary-web non-discovery cannot establish novelty.","source_ids":["S1","S2","S3","S4","S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Existing products and a standard demonstrate typed statuses, failure adapters, recovery policies, and audit semantics. Heterogeneous promise enforcement and downstream label preservation still require integration testing.","source_ids":["S1","S2","S3","S4"]},"adoption_authority_feasibility":{"score":3,"rationale":"NASA guidance supplies a credible decision-authority model, and a board-authorized shadow test is feasible. Production adoption requires board, data-owner, vendor, laboratory, privacy, and security approval.","source_ids":["S7"]},"evidence_readiness":{"score":4,"rationale":"A falsifiable shadow comparison can be run on at most 20 recorded or scripted cases without altering decisions, provided proprietary traces and evaluator owners are available.","source_ids":["S1","S2","S3","S4","S6"]},"safety_net_benefit":{"score":4,"rationale":"Explicit OUT_OF_PROMISE, UNRESOLVED, and ESCALATED states reduce false certainty and preserve board authority. Excessive or erroneous escalation could impede operations, so the benefit is not unconditional.","source_ids":["S2","S4","S7"]},"scalability":{"score":3,"rationale":"A shared contract schema and router can be reused, but every evaluator requires a maintained promise, adapter, authority rule, and change trigger; human and vendor interfaces may resist mechanical standardization.","source_ids":["S1","S2","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Prepare and run a non-decisional shadow study of at most 20 recorded or scripted evaluations; define route labels and comparators; review traces; and report failures and coverage gaps.","confidence":"MODERATE","assumptions":["Approximately 100-250 combined software, optimization, systems-engineering, and review-board labor hours.","A 2026 fully loaded internal rate of roughly $100-$180 per hour, inferred from BLS mean wages plus assumed benefits, overhead, and management.","Existing workflow logs or scripted cases are available without purchasing laboratory experiments or vendor simulation runs.","Security and legal review is limited to approving a contained shadow dataset."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build a one-workflow prototype with a contract schema, promise checker, four evaluator adapters, typed route engine, audit record, and shadow dashboard.","confidence":"MODERATE","assumptions":["Roughly 3-8 person-months across software, systems engineering, domain review, and testing.","Existing optimizer and gate systems expose integration points.","No major commercial licensing, laboratory instrumentation, or procurement changes are required.","The estimate excludes production accreditation and organization-wide rollout."],"source_ids":["S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Productionize one program's workflow, including security hardening, access controls, records integration, verification, governance approval, training, monitoring, and rollback exercises.","confidence":"LOW","assumptions":["Roughly 10-30 person-months plus security, governance, vendor coordination, and contingency.","Four evaluator classes can expose stable interfaces or documented manual adapters.","Production validation must show that labels persist through every optimizer, report, dashboard, export, and gate interface.","Costs of new laboratory tests, proprietary simulators, or major platform replacement are excluded."],"source_ids":["S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain contracts and adapters, review evaluator and policy changes, audit route records, handle incidents, retrain users, and revalidate gate-label propagation.","confidence":"LOW","assumptions":["Approximately 0.5-1.5 full-time-equivalent engineering and governance effort plus modest tooling.","Evaluator interfaces and gate policies change incrementally rather than being replaced annually.","External simulation, laboratory, and vendor usage fees remain separate operational costs."],"source_ids":["S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent product documentation, a standard, and primary studies directly document failed, missing, misleading, or inconsistent evaluations in simulation and human-guided optimization.","source_ids":["S1","S2","S3","S4","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"NASA's program decision authority or decision-making body is an identifiable, credible authorizer class, and NASA guidance explicitly requires sufficient technical evidence plus documented uncertainty, assumptions, and limitations. This does not establish adoption intent for the candidate.","source_ids":["S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim is contrastive and falsifiable on fixed traces: zero false-definitive gate verdicts under required non-results versus the current and dummy-value/NaN-style comparators, with no valid in-promise successes blocked. Its novelty is not established.","source_ids":["S1","S2","S3","S4","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A shadow study capped at 20 cases has specified evaluator strata, comparators, recorded outputs, pass conditions, and falsifiers and does not require production changes.","source_ids":["S1","S2","S3","S4","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"For the bounded shadow step, the review board retains authority, no live verdict changes, typed non-results cannot authorize action, original records are preserved, and the exercise can be disabled. Data-owner approval and controlled handling of vendor, laboratory, and human records are prerequisites; their absence stops the study rather than creating an unsafe route.","source_ids":["S2","S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands are tied to explicit labor, integration, governance, and exclusion assumptions and anchored to current official wage data. Confidence declines for production and recurring costs because no partner architecture or vendor pricing was available.","source_ids":["S8"]}},"next_evidence_step":"With design-review-board and data-owner approval, pre-register a non-decisional shadow comparison using at most 20 recorded or scripted candidate evaluations. Include at least one case from each evaluator class—simulator, laboratory, vendor, and human—and cover success plus every available instance of OUT_OF_PROMISE, delay, refusal, inconsistency, error, and no result. Compare (A) the workflow's original route and gate label, (B) a documented dummy-value/NaN-style solver-local comparator, and (C) the proposed typed contract and router. Record promise-check result, evidence provenance, evaluator state, retry or escalation route, downstream optimization label, shadow gate label, latency, and whether a valid in-promise success was blocked. The nominated problem is falsified if baseline A already preserves every required non-result as typed and non-definitive. The intervention is falsified if C produces any unqualified feasible, infeasible, or optimal verdict after a required non-result, blocks any valid in-promise success, loses a typed label downstream, or cannot define an enforceable promise for one of the four classes. Treat unobserved failure modes as coverage gaps, not passes; halt if shadow records affect a live gate or expose unauthorized data.","blocking_evidence":["No representative proprietary workflow traces were available to measure false-definitive verdict prevalence or compare routes.","No named organization has committed to adopt, fund, or authorize the proposed router.","Mechanical and semantically valid promises have not been demonstrated for laboratory, vendor, and human evaluators.","Downstream preservation of typed labels through an actual optimizer, reporting stack, and gate interface has not been tested.","Escalation load, false-blocking rate, latency effects, and operator usability are unmeasured.","Vendor terms, laboratory-data controls, human-evaluator privacy, security, and records-management requirements require partner-specific review.","Production and recurring cost bands lack partner architecture, procurement, and vendor-price data.","The bounded search cannot establish absence of deployed integrations, patents, unpublished practices, or additional prior art."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured. The eight-source bounded search found substantial prior art for external-analysis failure routing, typed statuses, undefined outputs, NaN propagation, hidden-constraint handling, inconsistent human judgments, and decision provenance. It did not find one implementation combining mechanically checked promises for simulator, laboratory, vendor, and human evaluators with preservation of UNRESOLVED or ESCALATED through the final multidisciplinary gate. That non-discovery defines only the residual search boundary and is not evidence of worldwide novelty.","arm":"RETRIEVAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Secure written design-review-board, workflow-owner, and data-owner authorization for a non-decisional shadow study.","Obtain or script at most 20 auditable cases spanning all four evaluator classes and available non-result modes.","Pre-register the current-workflow, dummy-value/NaN, and typed-router comparators and the zero-false-definitive/no-valid-block falsifiers.","Demonstrate enforceable promises and typed adapters for laboratory, vendor, and human evaluators, or narrow the claim to the classes that can be enforced.","Measure downstream label preservation, false blocking, escalation load, latency, coverage gaps, and operator usability.","Complete partner-specific privacy, vendor-terms, laboratory-data, security, records, and production-authority review.","Replace broad startup, launch, and recurring cost bands with partner architecture and procurement estimates.","If empirical results remain favorable, conduct a separate prior-art differentiation and freedom-to-operate study without treating ordinary-web absence as novelty."],"reason":"Web evidence verifies the problem, a credible authorizer class, substantial constituent prior art, and technical plausibility, but it cannot establish the candidate's incremental performance or heterogeneous integration. The decisive remaining evidence requires proprietary workflow traces, stakeholder authorization, and shadow execution; therefore further bounded web search is not the controlling next step."}}