{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"representation_independent_interface_contract__human_computer_interaction:P2:v0","cell_id":"representation_independent_interface_contract__human_computer_interaction","search_queries":["site:gov.uk service manual analytics redesign events consistent event taxonomy user interactions","software telemetry instrumentation changes UI redesign longitudinal analytics data quality study","OpenTelemetry semantic conventions schema evolution stability official","xAPI specification profiles verbs activities conformance official","CHI user interaction logging instrumentation data quality interface changes analytics study","\"redesign\" \"event tracking\" analytics instrumentation usability","site:gov.uk \"event taxonomy\" analytics interaction","usability telemetry longitudinal redesign instrumentation events research","site:docs.publishing.service.gov.uk/analytics GA4 event schema consistent maintain data quality universal event structure","site:snowplow.io/docs data contracts tracking plans schema evolution events official","site:adlnet.gov xAPI 2.0 specification profiles conformance official","site:iso.org ISO 9241-11 usability effectiveness efficiency satisfaction context use","1EdTech Caliper Analytics specification metric profiles event model conformance official","ADL xAPI profile specification 2.0 official verbs patterns validation","W3C privacy principles data minimization user interaction telemetry official","NIST Privacy Framework data minimization telemetry user data official"],"sources":[{"source_id":"S1","title":"A Reference Data Model for Process-Related User Interaction Logs","publisher":"Luka Abb and Jana-Rebecca Rehse / arXiv","url":"https://arxiv.org/abs/2207.12054","source_class":"PRIMARY_RESEARCH","publication_date":"2022-07-25","accessed_at":"2026-08-03","claims_supported":["UI logs commonly encode low-level actions, target elements, UI hierarchy, and sometimes raw input values.","Research and industry UI-log formats vary substantially, obstructing integration with downstream analytics.","Mapping low-level interactions to conceptual tasks is necessary for meaningful analysis.","An application-independent UI-log reference model and standardized exchange format are technically feasible."]},{"source_id":"S2","title":"Telemetry Schemas","publisher":"OpenTelemetry","url":"https://opentelemetry.io/docs/specs/otel/schemas/","source_class":"STANDARD","publication_date":"n.d.; stable specification accessed 2026-08-03","accessed_at":"2026-08-03","claims_supported":["Implicit assumptions about telemetry shape couple producers, conventions, and consumers.","Versioned schemas and explicit transformations permit independently evolving telemetry producers and consumers.","Non-convertible schema changes must be treated as breaking changes.","The OpenTelemetry schema mechanism does not fully define telemetry values or behavioral semantics."]},{"source_id":"S3","title":"Defining the data to collect with Tracking Plans","publisher":"Snowplow","url":"https://docs.snowplow.io/docs/event-studio/tracking-plans/","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-07-23","accessed_at":"2026-08-03","claims_supported":["Tracking plans already provide owned event specifications combining technical structure and business context.","Event specifications can govern instrumentation, generate type-safe code, monitor collection, and synchronize downstream models.","Product and analytics teams are identifiable adopters of governed event contracts."]},{"source_id":"S4","title":"Caliper Analytics 1.2 Specification","publisher":"1EdTech Consortium (formerly IMS Global)","url":"https://www.imsglobal.org/spec/caliper/v1p2/","source_class":"STANDARD","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["Caliper already specifies structured semantic interaction events, controlled actions, profiles, contextual entities, and conformance requirements.","Semantic interaction records can be produced consistently by multiple applications for downstream decision-making.","Caliper is domain-specific to learning activity and does not establish redesign-invariance for general usability evidence."]},{"source_id":"S5","title":"Privacy Principles","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/privacy-principles/","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025","accessed_at":"2026-08-03","claims_supported":["Behavioral information flows can create privacy risks even when data is not obviously identifying.","Data collection and transfer should be restricted to what is necessary for users' goals or aligned with their wishes and interests.","Data minimization reduces disclosure and misuse risk, supporting a bounded telemetry contract."]},{"source_id":"S6","title":"ISO 9241-11:2018 — Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts","publisher":"International Organization for Standardization","url":"https://www.iso.org/standard/63500.html","source_class":"STANDARD","publication_date":"2018-03","accessed_at":"2026-08-03","claims_supported":["Usability is an outcome of use in a specified context, not merely an event-schema property.","The standard supplies concepts but not a specific instrumentation or evaluation method.","Contract conformance alone cannot establish the validity of a usability construct or metric."]},{"source_id":"S7","title":"govuk_publishing_components: Data schemas","publisher":"Government Digital Service / GOV.UK","url":"https://docs.publishing.service.gov.uk/repos/govuk_publishing_components/analytics-ga4/schemas.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2024-03-27","accessed_at":"2026-08-03","claims_supported":["GOV.UK Publishing is an identifiable adopter of centrally governed analytics schemas.","Central schemas let GOV.UK alter structure or add attributes without updating hundreds of element-level data attributes.","Schema validation and ignoring undeclared fields demonstrate practical boundary enforcement, although the page warns it may be outdated."]},{"source_id":"S8","title":"Analytics Best Practice Guide: Going Through a Redesign","publisher":"Seer Interactive","url":"https://www.seerinteractive.com/insights/analytics-best-practices-redesign","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"2018-09-17","accessed_at":"2026-08-03","claims_supported":["Website redesigns materially affect analytics configuration and continuity.","Historical comparison is feasible only when objectives, actions, and functionality remain sufficiently comparable.","Changed sections, goals, event structures, or user experience require re-evaluation rather than automatic longitudinal comparison."]}],"problem_evidence":{"support":"MODERATE","rationale":"The problem is externally visible: primary research documents incompatible, ad hoc UI-log concepts and preprocessing burdens; OpenTelemetry documents producer-consumer coupling under telemetry evolution; and redesign guidance identifies breaks in direct comparison when interface structure or desired actions change. GOV.UK's central-schema design further demonstrates the maintenance burden of element-level instrumentation. Evidence does not measure prevalence, labor cost, or decision error for customer-support consoles specifically.","source_ids":["S1","S2","S7","S8"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"GOV.UK Publishing is an identifiable analogous adopter with a central analytics schema intended to avoid distributed updates, while Snowplow explicitly assigns tracking-plan owners and describes cross-team governance needs. These sources establish credible analytics, instrumentation, and governance owners, but no organization was found requesting the proposal's exact privacy-bounded task-state contract or committing staff or funds.","source_ids":["S3","S7","S8"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"1EdTech Caliper Analytics","similarity":"Defines semantic activity events, controlled action vocabularies, domain profiles, contextual entities, and conformance requirements so multiple applications can produce comparable interaction records.","remaining_difference":"Caliper targets learning analytics and does not require presentation-mutation invariance, a privacy-exclusion oracle, or a redesign-specific admission gate for longitudinal usability evidence.","source_ids":["S4"]},{"name":"Snowplow Tracking Plans and event specifications","similarity":"Provides owned event contracts combining business meaning, triggers, schemas, implementation instructions, validation, generated bindings, observability, and synchronized downstream models.","remaining_difference":"The documentation does not provide a task-evidence state machine, operation-sequence laws, independently implemented redesign adapters, or metamorphic tests asserting that relabeling and layout changes preserve evidence.","source_ids":["S3"]},{"name":"OpenTelemetry Telemetry Schemas","similarity":"Separates versioned telemetry schemas from producers and consumers, identifies breaking changes, and specifies transformations between compatible versions.","remaining_difference":"It governs telemetry-schema evolution rather than the semantic equivalence of human task evidence; it expressly does not fully specify values or behavioral semantics.","source_ids":["S2"]},{"name":"Reference Data Model for Process-Related User Interaction Logs","similarity":"Standardizes heterogeneous UI logs, supports multiple abstraction levels, adds task context, and demonstrates mapping interactions to higher-level tasks independently of recording approach.","remaining_difference":"Its core model retains actions, target elements, UI hierarchy, and potentially input values; it lacks longitudinal substitution gates, privacy exclusions, state-transition rejection semantics, and presentation-mutation conformance testing.","source_ids":["S1"]},{"name":"GOV.UK centralized GA4 data schemas","similarity":"Uses a central schema to validate analytics records, ignore undeclared attributes, and change structure without updating hundreds of local element annotations.","remaining_difference":"It remains an event/data-layer schema tied to interface interactions and does not claim task-semantic equivalence across redesigns.","source_ids":["S7"]}],"distinctive_claim_remaining":"For an unchanged, explicitly versioned task, two independently authored release adapters can map structurally different interface traces into the same privacy-bounded task-evidence state while rejecting seeded transition and metadata violations; presentation-only mutations must not change that state, whereas genuine task-semantic changes must force a version boundary. This is contrastive and falsifiable but untested.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Machine-readable schemas, generated bindings, central validation, semantic event models, profiles, version transformations, and conformance processes are established and technically implementable. An offline two-adapter study therefore needs no novel infrastructure. The difficult step is organizational and semantic: researchers must predefine task meaning, resolve ambiguous mappings, prevent trend-preserving adapter tuning, and distinguish presentation-dependent usability phenomena from contract-level evidence. Production collection would additionally require privacy/workforce-data authority and purpose limitations; the proposed synthetic-only first step avoids that immediate safety stop.","source_ids":["S1","S2","S3","S4","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Avoiding false longitudinal discontinuities and reducing presentation-linked data collection could materially improve decisions, but target prevalence, labor burden, and decision consequences are unmeasured.","source_ids":["S1","S8"]},"stakeholder_pull":{"score":3,"rationale":"Analytics teams demonstrably use governed schemas and seek continuity through redesign, but there is no confirmed adopter requesting this exact intervention.","source_ids":["S3","S7","S8"]},"incremental_advantage":{"score":3,"rationale":"State laws, negative controls, privacy exclusions, and presentation-mutation tests add a meaningful verification layer beyond naming conventions and warehouse schemas, though the advantage has not been measured against existing tracking-plan practice.","source_ids":["S2","S3","S4"]},"distinctiveness_plausibility":{"score":2,"rationale":"The proposal combines several established mechanisms and substantially collides with Caliper, tracking plans, telemetry schemas, and task-context UI-log models. The precise redesign-invariance test remains differentiated, not world-novelty verified.","source_ids":["S1","S2","S3","S4"]},"technical_implementability":{"score":4,"rationale":"Schemas, adapters, state machines, generated bindings, validation, and conformance suites are mature techniques; semantic ambiguity rather than infrastructure is the main technical risk.","source_ids":["S2","S3","S4","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"Instrumentation, usability-research, and privacy owners are identifiable and a synthetic sandbox is within ordinary internal authority. Production use involving worker interaction records requires additional governance approval and purpose restrictions.","source_ids":["S3","S5","S7"]},"evidence_readiness":{"score":4,"rationale":"A small synthetic mutation study can directly test the remaining claim without production telemetry or human-subject exposure. No such test results currently exist.","source_ids":["S2","S3","S4"]},"safety_net_benefit":{"score":4,"rationale":"Field allowlists, rejection of raw content and identifiers, synthetic first evidence, explicit invalid states, and rollback rules create useful safeguards against collection creep and false comparability, although downstream misuse remains possible.","source_ids":["S5","S6"]},"scalability":{"score":3,"rationale":"Reusable schemas and conformance tooling can scale across releases, but each distinct task requires semantic modeling, adapter maintenance, version adjudication, and leakage review.","source_ids":["S1","S2","S3","S4"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"One offline task, ten predeclared synthetic journeys, two independently written adapters, presentation mutations, seeded invalid transitions and prohibited metadata, and a short privacy/research review.","confidence":"MODERATE","assumptions":["Approximately 6-10 person-weeks across an HCI researcher, two instrumentation engineers, and part-time privacy review.","Existing prototype traces, schema tooling, and test infrastructure are available.","No production telemetry, participant recruitment, procurement, or external certification."],"source_ids":["S2","S3","S4","S5"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Production-grade contract and adapter toolkit for one support workflow, CI conformance gate, version registry, leakage checks, documentation, and governance workflow, excluding production launch.","confidence":"LOW","assumptions":["Approximately 3-6 loaded person-months.","Existing analytics transport and warehouse remain in place.","One task family and two interface releases are covered."],"source_ids":["S2","S3","S7"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Privacy-approved production integration across several support-console workflows, dual-running with current telemetry, dashboard migration, access controls, incident/rollback procedures, training, and release governance.","confidence":"LOW","assumptions":["Approximately 8-18 loaded person-months across engineering, research, analytics, security/privacy, and program management.","No replacement of the underlying warehouse or analytics vendor.","Employee consultation, legal review, and security assessment are required but no litigation or major procurement is assumed."],"source_ids":["S3","S5","S7"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Contract stewardship, adapter updates, conformance CI maintenance, periodic leakage audits, task-version adjudication, documentation, and support for several releases per year.","confidence":"LOW","assumptions":["Approximately 2-5 loaded person-months annually plus modest infrastructure cost.","The number of governed task families remains bounded.","No continuous manual review of individual worker sessions."],"source_ids":["S2","S3","S5"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent primary, standards, government, and practitioner sources show heterogeneous UI logs, producer-consumer telemetry coupling, distributed schema maintenance, and redesign-related comparison discontinuities.","source_ids":["S1","S2","S7","S8"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"GOV.UK Publishing is an identifiable analogous adopter of centrally governed analytics schemas, and Snowplow documents explicit tracking-plan ownership and cross-team governance. This verifies credible authority, not commitment to this proposal.","source_ids":["S3","S7"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim is narrower than identified prior art and can be falsified by disagreement between independent adapters, sensitivity to presentation-only mutations, failure to reject seeded violations, or failure to version genuine task changes.","source_ids":["S1","S2","S3","S4"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-task, ten-journey, synthetic two-adapter mutation study is time-, data-, and scope-bounded and has explicit incumbent and contract comparators plus predetermined falsifiers.","source_ids":["S2","S3","S4"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The first step is offline and synthetic, prohibits production collection and identifiers, and can be stopped by deleting sandbox records. Production remains contingent on privacy and workforce-data authorization.","source_ids":["S5","S6"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four bands are tied to explicit task counts, person-month assumptions, existing infrastructure, and exclusions. Monetary confidence remains moderate-to-low because no target organization's loaded labor rates or integration architecture were available.","source_ids":["S2","S3","S7"]}},"next_evidence_step":"Run a four-week offline mutation study on one support task. Before seeing adapter outputs, a usability researcher and instrumentation owner specify ten contract-level journeys and the exact accepted state for each. Two engineers independently implement adapters for an incumbent synthetic trace and a structurally different prototype trace. Compare (A) the proposed state-machine contract against (B) the current locally named-event plus manual warehouse-map baseline. Apply at least three presentation-only mutations per journey—control relabeling, layout/navigation reordering, and equivalent keyboard activation—and seed orphan outcomes, duplicate closure, post-closure events, undeclared semantic relabeling, and prohibited raw text or identifiers. Pass only if all valid paired journeys produce exactly the same abstract state under the proposal, all seeded violations are rejected without state mutation, no excluded field reaches output, and reviewers correctly designate every predeclared genuine task change as a new version. Falsify or revise the intervention if any presentation-only mutation changes abstract evidence, either independent adapter disagrees on a predeclared journey, a seeded violation passes, agreement requires interface structure or prohibited content, or the baseline performs equally well with no greater review effort. Do not collect worker data or infer usability, confusion, workload, or satisfaction.","blocking_evidence":["No executed two-adapter mutation study establishes representation invariance or negative-control sensitivity.","No target-console audit measures the frequency, labor cost, or decision consequences of redesign-induced telemetry discontinuities.","No named adopter has committed staff, funding, or production authority for this exact task-evidence contract.","No privacy or workforce-data authority has approved production metadata, purposes, retention, access, or separation from performance management.","No evidence shows that contract-derived measures validly estimate usability constructs or improve realized decisions.","The comparative maintenance burden versus a well-governed tracking plan and semantic warehouse layer is unknown."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes substantial adjacent and colliding practice, not world novelty. Patentability, freedom to operate, exhaustive product or literature coverage, market size, prevalence, realized impact, metric validity, and legal compliance in any particular jurisdiction remain unmeasured. The only surviving novelty-like boundary is the specific combination of a privacy-bounded task-evidence state machine, independently implemented release adapters, presentation-mutation conformance tests, and an admission rule for longitudinal usability evidence.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Execute the pre-registered two-adapter synthetic mutation study and report every journey, mutation, seeded violation, and disagreement.","Benchmark correctness and reviewer effort against a conventional tracking plan plus release-specific warehouse mapping.","Obtain written usability-research, instrumentation, and privacy/workforce-data authority decisions before any production telemetry consideration.","Measure redesign-discontinuity prevalence and remediation effort in one target console using existing records only.","Demonstrate that agreement does not require DOM paths, labels, screen order, raw content, coordinates, or stable personal identifiers.","Maintain the explicit boundary that conformance does not validate a usability construct or justify employee-performance use."],"reason":"Bounded web research verifies the generic problem, credible analogous adopters, technical feasibility, privacy constraints, and substantial prior-art collision. It cannot establish the surviving causal claim: that independent adapters remain semantically equivalent under redesign mutations while outperforming existing governed event-contract practice. That evidence requires implementation and live execution of an offline empirical test, so further web research is not the controlling next step."},"proposal_index":2}