{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"representation_independent_interface_contract__futurism_foresight:P4:v0","cell_id":"representation_independent_interface_contract__futurism_foresight","search_queries":["site:gov.uk Department for Transport futures toolkit uncertainty strategic foresight transport forecasting","site:metaculus.com/help question resolution criteria scoring predictions API","forecasting platform question resolution criteria probability updates scoring standard research reproducibility","Good Judgment Open forecasting question resolution rules scoring API","forecasting questions common data schema API resolution scoring update history open source","forecasting platform interoperability standard question schema resolution criteria probabilities","site:github.com forecasting question schema resolution scoring API prediction history","site:arxiv.org forecasting platform reproducibility data schema question resolution scoring","site:gov.uk probabilistic forecasting government evaluation Brier score questions forecasting platform","site:gov.uk transport strategic foresight probabilistic forecasts forecasting questions","government forecasting platform probability questions scoring resolution strategic foresight agency","national transport agency strategic foresight forecast probability judgments platform"],"sources":[{"source_id":"S1","title":"TAG uncertainty toolkit","publisher":"UK Department for Transport","url":"https://www.gov.uk/government/publications/tag-uncertainty-toolkit","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2021-05-19; updated 2026-05-28","accessed_at":"2026-08-03","claims_supported":["A national transport department maintains guidance for analysing and presenting uncertainty in future transport demand.","Transport uncertainty analysis is intended to inform investment decisions and business cases.","The department is an identifiable potential authorizer for methods used in transport appraisal."]},{"source_id":"S2","title":"Appraisal, Modelling and Evaluation Strategy","publisher":"UK Department for Transport","url":"https://assets.publishing.service.gov.uk/media/6a18628c65bc5f798327f4d1/dft-appraisal-modelling-evaluation-strategy-for-transport.pdf","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05","accessed_at":"2026-08-03","claims_supported":["Long-lived transport decisions face substantial uncertainty from behavioural, technological, climate and decarbonisation trends.","The department reports a strategic need to refresh uncertainty guidance because current approaches can be difficult for decision-makers to interpret.","The department is modernizing analysis through automation, improved interoperability between tools, transparency and accessibility."]},{"source_id":"S3","title":"Metaculus FAQ","publisher":"Metaculus","url":"https://www.metaculus.com/faq/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["An operating forecasting platform already represents mutually exclusive and exhaustive outcomes whose probabilities sum to 100%.","Metaculus implements resolution, annulment and scoring behavior, including unscored annulled questions.","Superficially similar question groups and multiple-choice questions have materially different probability and scoring semantics."]},{"source_id":"S4","title":"Metaculus Question Approval Checklist","publisher":"Metaculus","url":"https://www.metaculus.com/help/question-checklist/","source_class":"OFFICIAL_GUIDANCE","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Forecast resolution requires explicit authority and carefully specified sources, dates and criteria.","Metaculus describes itself as both a tracker of current information and a record and evaluation of past predictions.","Ambiguous wording, relative dates and dependent criteria create durable interpretation problems for long-horizon questions."]},{"source_id":"S5","title":"Scores FAQ","publisher":"Metaculus","url":"https://www.metaculus.com/help/scores-faq/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["Metaculus scores timestamped forecast updates according to how long each prediction was active.","Early resolution, scheduled closure and score truncation affect computed scores.","Metaculus has replaced scoring methods over time while retaining legacy scores, demonstrating the importance of scoring-rule version and lifecycle semantics."]},{"source_id":"S6","title":"Good Judgment Open FAQ","publisher":"Good Judgment Open","url":"https://www.gjopen.com/faq?source=post_page---------------------------","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"resolution guidance updated 2024-06-03","accessed_at":"2026-08-03","claims_supported":["Good Judgment Open uses platform-specific rules for evidentiary resolution, voiding, retroactive closure, time zones, rounding and data revisions.","It uses ordinary and ordered-categorical scoring rules and does not permit forecast withdrawal or deletion.","Its rules differ visibly from Metaculus rules, supporting the risk that identical displayed probabilities need not imply equivalent lifecycle or score semantics."]},{"source_id":"S7","title":"Cultivate Forecasts","publisher":"Cultivate Labs","url":"https://www.cultivatelabs.com/forecasts","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"n.d.","accessed_at":"2026-08-03","claims_supported":["A commercial product already provides question authoring, probability updates, resolution, Brier-style scoring, analytics and exportable data as a complete lifecycle.","Cultivate reports deployment of a government-wide forecasting program for the UK Cabinet Office.","Existing product functionality makes the proposed component technically plausible but also creates substantial prior-art overlap."]},{"source_id":"S8","title":"Alignment Problems With Current Forecasting Platforms","publisher":"arXiv","url":"https://arxiv.org/abs/2106.11248","source_class":"PRIMARY_RESEARCH","publication_date":"2021-06-21","accessed_at":"2026-08-03","claims_supported":["The authors compare scoring and incentive behavior across Metaculus, Good Judgment Open, CSET-Foretell and Cultivate-based platforms.","They identify reward-specification problems and show that implementation and scoring choices can change incentives and evaluation.","The paper supports the consequence of scoring semantics, but does not study transport-agency migrations or representation-independent conformance contracts."]}],"problem_evidence":{"support":"MODERATE","rationale":"The problem class is visible: DfT identifies consequential future uncertainty and a need for interoperable, transparent analytical tools, while first-party platform rules show that closure, update, resolution, time-zone, rounding, cancellation and scoring semantics differ materially. Metaculus has also changed scoring methods while retaining legacy scores. However, no direct source documents a transport agency whose forecast scores changed during an export, platform migration or reanalysis; that central prevalence claim remains unverified.","source_ids":["S1","S2","S3","S5","S6","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"The UK Department for Transport is an identifiable authorizer with published responsibility for uncertainty guidance and an expressed interest in tool interoperability, automation and transparency. Cultivate reports that the UK Cabinet Office operates a government-wide forecasting program, establishing public-sector adoption of resolvable probability forecasting. Neither source expresses demand for a platform-neutral behavioral contract or conformance suite specifically, so pull is adjacent rather than direct.","source_ids":["S1","S2","S7"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Cultivate Forecasts lifecycle","similarity":"Provides question authoring, probability updates, resolution, Brier-style scoring, auditability, analytics and exportable data for enterprise and government programs.","remaining_difference":"The public description does not establish a representation-independent state-machine specification, portable versioned contract or common black-box conformance suite for competing implementations.","source_ids":["S7"]},{"name":"Metaculus question, resolution and scoring model","similarity":"Implements exhaustive outcome sets, probability constraints, timestamped updates, resolution authority, annulment, score truncation and legacy scoring methods.","remaining_difference":"It is a platform-specific operational policy rather than a cross-platform substitution contract; no reusable conformance oracle was found.","source_ids":["S3","S4","S5"]},{"name":"Good Judgment Open forecasting lifecycle","similarity":"Defines update retention, closure, resolution evidence, voiding, rounding, time zones and multiple scoring treatments.","remaining_difference":"Its rules are platform-specific and sometimes deliberately preserve resolver discretion; they do not establish equivalence across document, ledger and platform representations.","source_ids":["S6"]},{"name":"Cross-platform scoring-alignment analysis","similarity":"Treats scoring-rule and implementation choices as consequential specifications and compares several forecasting platforms.","remaining_difference":"It evaluates incentives rather than defining a lifecycle ADT, migration semantics or executable substitution tests.","source_ids":["S8"]}],"distinctive_claim_remaining":"For a predeclared corpus of categorical forecast histories, a versioned, platform-neutral state-machine contract plus one reusable black-box and differential conformance suite will make independently implemented document, ledger and scoring adapters return identical contract-level active forecasts, typed errors, terminal states and scores, while exposing any representation-dependent timezone, precision, ordering or identity leakage. The claim is falsified if two implementations pass the suite yet disagree on any predeclared observable, or if an ordinary approved claim cannot be represented without implementation-specific fields.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Existing platforms already implement nearly all required operations and demonstrate technical feasibility of timestamped updates, finite outcome sets, resolution, cancellation or voiding, scoring, analytics and export. An offline event-ledger reference model and adapters are ordinary software-engineering work. Feasibility is not yet established for the agency's actual data formats, undocumented scoring state, identity controls, approved resolution procedures or integration boundaries. Human resolution judgment also cannot safely be replaced by conformance testing.","source_ids":["S3","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Stable interpretation and evaluation could improve institutional learning and prevent misleading retrospective score changes, but the prevalence and decision consequences of migration-induced divergence are not measured.","source_ids":["S1","S2","S5","S6","S8"]},"stakeholder_pull":{"score":3,"rationale":"Government transport uncertainty work and a Cabinet Office forecasting program identify plausible authorizers and users, but neither has requested this contract.","source_ids":["S1","S2","S7"]},"incremental_advantage":{"score":3,"rationale":"Portable behavioral equivalence testing could improve on schemas, platform mandates and manual reconciliation, but no comparative evidence yet shows lower error or migration effort.","source_ids":["S3","S5","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"Cross-implementation conformance is a specific remaining distinction, although most lifecycle elements are already established practice and the search did not establish world novelty.","source_ids":["S3","S4","S5","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"Existing products demonstrate the constituent state, update, resolution and scoring functions; the first offline test is technically bounded.","source_ids":["S3","S5","S6","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"A transport department can govern analytical guidance and internal tools, but resolution committees, data owners, platform maintainers and personnel-evaluation authorities would need separate approvals.","source_ids":["S1","S2","S7"]},"evidence_readiness":{"score":3,"rationale":"Public rules support a synthetic comparator suite immediately, but decision-relevant validation requires access to closed claims, exports and scoring code.","source_ids":["S3","S4","S5","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"A read-only offline pilot using synthetic or de-identified closed claims is reversible and can detect discrepancies without changing authoritative forecasts or scores.","source_ids":["S4","S5","S6"]},"scalability":{"score":4,"rationale":"A shared specification and parameterized suite can be reused across adapters and scoring versions, although each legacy representation still requires mapping and stewardship.","source_ids":["S3","S5","S6","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Four to six weeks to specify a minimal categorical-claim state model, build a synthetic corpus, implement one simple reference model and compare two offline adapters.","confidence":"MODERATE","assumptions":["Approximately 160-320 combined engineering, domain-review and test-design hours.","No connection to authoritative systems and no identity-bearing data.","Resource-equivalent estimate, not a vendor quotation."],"source_ids":["S3","S4","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Versioned specification, production-quality conformance harness, two or three read-only adapters, governance review, privacy review and migration documentation.","confidence":"LOW","assumptions":["Three to six staff-months across software engineering, data stewardship, forecasting methodology and governance.","Existing systems expose usable exports or read APIs.","No replacement of the authoritative platform."],"source_ids":["S2","S5","S6","S7"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production integrations, security and access controls, validation of historical records, workflow training, resolution-committee procedures, monitoring and staged rollout.","confidence":"LOW","assumptions":["Multiple legacy representations and at least one production forecasting platform.","Independent security, privacy and records-management reviews are required.","Historical discrepancies are triaged rather than silently rewritten."],"source_ids":["S2","S7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Contract stewardship, adapter maintenance, scoring-version governance, regression testing, leakage audits and user support.","confidence":"LOW","assumptions":["Approximately 0.5-1.5 full-time-equivalent staff plus infrastructure and periodic review.","A modest internal program rather than a national shared service.","Major platform migrations are excluded."],"source_ids":["S2","S5","S7"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Consequential uncertainty, demand for analytical interoperability and materially different platform lifecycle and scoring rules are externally visible, although transport-specific migration failures are not yet measured.","source_ids":["S1","S2","S5","S6","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The UK Department for Transport is a credible analytical-method authorizer, and public-sector forecasting adoption is evidenced by the reported UK Cabinet Office program.","source_ids":["S1","S2","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim concerns cross-implementation equality of specified observables under one versioned contract and has explicit failure conditions.","source_ids":["S3","S5","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A read-only, time-boxed comparison of a reference ledger and two adapters over a predeclared synthetic and closed-claim corpus has defined comparators and falsifiers.","source_ids":["S3","S4","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first test can avoid live writes, personal scoring and operational decisions. Resolution and cancellation authority must remain with the designated committee, and any identity leakage or need to alter source records is a halt condition.","source_ids":["S4","S6"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Resource scopes and broad bands can be stated, but no system inventory, labor rate, procurement quote, integration count or agency security requirement was available.","source_ids":["S2","S7"]}},"next_evidence_step":"Run a six-week partnered, read-only pilot. Before access, freeze 40-60 synthetic histories covering valid and invalid distributions, multiple updates, both sides of closure boundaries, retroactive closure, disputes, cancellations, time zones and precision edges; then add 20-30 de-identified closed agency claims without altering them. Compare (A) the incumbent export-plus-scoring script, (B) an independent event-ledger reference model and (C) a document-backed adapter. Measure equality of active forecasts, accepted and rejected transitions, terminal states and scores, plus implementation effort and leakage findings. Falsify the intervention if two implementations pass the full suite but differ on a predeclared observable, if at least one ordinary approved claim is inexpressible without a hidden field, or if scoring requires undocumented state. Falsify the stated problem locally if all three implementations agree on every difficult history and current records contain all required semantics without undocumented defaults.","blocking_evidence":["No direct evidence that a transport agency currently maintains the same forecast across multiple representations or has experienced migration-induced score drift.","No adopter has expressed demand for a platform-neutral behavioral contract or agreed to own its stewardship.","No public comparison shows that existing platform exports omit the semantics needed to reconstruct active forecasts and scores.","No agency system inventory or rate basis supports narrower deployment and recurring-cost estimates.","The proposed benefit over established platform lifecycle controls and careful versioned exports has not been measured.","Whether approved strategic forecasts can be reduced to finite, stable outcomes without losing decision-relevant context requires domain review and real claim examples."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"World novelty, patentability, freedom to operate, market size and realized impact were not measured. The search only establishes that constituent lifecycle functions are established and that a cross-implementation behavioral contract with a reusable conformance oracle was not found in the eight reviewed sources; absence from this bounded search is not evidence of novelty.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a transport-agency partner and written authorization for read-only use of de-identified closed claims and exports.","Pre-register the abstract lifecycle, comparator implementations, equality tolerances, leakage channels and falsifiers before examining agency discrepancies.","Quantify baseline divergence across the incumbent platform, exports and independent scoring scripts, including closure, timezone, precision and cancellation cases.","Demonstrate that the contract represents ordinary approved claims without suppressing material qualifications or displacing resolution authority.","Produce an evidence-based integration inventory and resource estimate before any production decision."],"reason":"Bounded web research established a credible problem class, adopter class, substantial prior-art overlap and a falsifiable remaining claim. It cannot establish transport-agency prevalence, expressibility, comparative advantage or actual integration cost. Those decisive questions require proprietary records, domain review and live offline testing, so further desk research alone cannot complete the evaluation."},"proposal_index":4}