{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__human_computer_interaction:P2:v0","cell_id":"predictive_residual_processing__human_computer_interaction","search_queries":["site:learn.microsoft.com clarity session recordings rage clicks dead clicks quick backs documentation","site:help.hotjar.com session replay frustration score filters rage clicks u-turns","site:help.fullstory.com session replay frustration signals funnels official","site:docs.logrocket.com Galileo AI session replay issue detection official","site:help.fullstory.com frustration signals session replay rage clicks official help","site:help.fullstory.com session replay identify user struggle funnels official","site:docs.logrocket.com Galileo AI issue detection session replay official docs","site:logrocket.com Galileo AI watches sessions surfaces severe issues","academic study user session replay usability anomaly detection interaction sequences task flow prediction UX","CHI paper session replay automatic usability issue detection interaction logs machine learning","research predicting next user action interface task model anomaly detection usability evaluation","systematic review automatic usability evaluation clickstream session replay user interaction anomalies","site:ico.org.uk session replay website privacy consent data minimisation user testing recordings","site:edpb.europa.eu guidelines consent website tracking recording data minimisation","site:w3.org privacy principles data minimization web interaction recording accessibility","session replay scripts privacy study sensitive data official academic paper","2026 UX researcher salary BLS official web developers digital interface designers wage","site:bls.gov occupational employment UX researcher human factors salary 2025","site:logrocket.com pricing session replay enterprise official","site:hotjar.com pricing session replay official","\"AXNav: Replaying Accessibility Tests from Natural Language\" PDF","\"AXNav\" Taeb Swearngin PDF","site:microsoft.com/research AXNav accessibility tests natural language"],"sources":[{"source_id":"S1","title":"Use Cases for Filtering Recordings","publisher":"Hotjar","url":"https://help.hotjar.com/hc/en-us/articles/36820019361041-Use-Cases-for-Filtering-Recordings","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["Session-replay users filter recordings to identify broken elements, confusion, frustration, and drop-off.","Existing practice detects predefined signals including rage clicks, U-turns, errors, exit pages, and entered-text events.","Filters narrow recordings to relevant sessions, establishing both stakeholder demand and a rule-based comparator."]},{"source_id":"S2","title":"Recordings overview","publisher":"Microsoft","url":"https://learn.microsoft.com/en-us/clarity/session-recordings/recordings-overview","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-05-21","accessed_at":"2026-08-03","claims_supported":["Clarity reconstructs sessions from captured HTML and user actions rather than literal video.","The product supports full-journey replay, event timelines, filters, segments, skip-inactivity, and AI summaries.","Current infrastructure demonstrates the feasibility of capturing and reconstructing chronological interaction traces."]},{"source_id":"S3","title":"Rage Clicks, Error Clicks, Dead Clicks, and Thrashed Cursor | Frustration Signals","publisher":"Fullstory","url":"https://help.fullstory.com/hc/en-us/articles/360020624154-Rage-Clicks-Error-Clicks-Dead-Clicks-and-Thrashed-Cursor-Frustration-Signals","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["Fullstory already surfaces session-replay moments using rage-click, error-click, dead-click, and thrashed-cursor signals.","The documentation acknowledges that intended repeated interactions can trigger rage-click heuristics, supporting the candidate's concern about legitimate alternative behavior and false positives.","Fullstory supports search refinement and muting, establishing configurable heuristic prioritization as prior art."]},{"source_id":"S4","title":"Introducing Galileo AI Summer ’24: Session summaries and severe mobile issues","publisher":"LogRocket","url":"https://blog.logrocket.com/product-management/galileo-ai-summer-24/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2024-06-25","accessed_at":"2026-08-03","claims_supported":["LogRocket states that teams manually tag events and watch hundreds of sessions, and that finding useful moments can consume hours or days.","Galileo watches sessions, identifies struggle and behavioral patterns, summarizes individual and grouped sessions, and links reviewers to problem areas.","Afp Capital is identified as an early adopter, and its IT-services manager expressed interest in using behavioral summaries to support customer-facing work.","Galileo substantially overlaps the candidate's attention-allocation and issue-surfacing function."]},{"source_id":"S5","title":"AXNav: Replaying Accessibility Tests from Natural Language","publisher":"arXiv; subsequently published at ACM CHI 2024","url":"https://arxiv.org/abs/2310.02424","source_class":"PRIMARY_RESEARCH","publication_date":"2023-10-03","accessed_at":"2026-08-03","claims_supported":["Manual accessibility testing is described as tedious, broad in scope, and difficult to schedule.","AXNav produces chaptered, navigable replay videos and flags potential accessibility issues with heuristics.","A study with 10 accessibility QA professionals found the system useful in their current work, supporting a credible human-review workflow analogue.","The work demonstrates technical feasibility for structured replay and attention-directing annotations but does not test predictive residual compression of natural user sessions."]},{"source_id":"S6","title":"Predicting Next Actions and Latent Intents during Text Formatting","publisher":"Collaborative Artificial Intelligence; CHI Workshop on Computational Approaches for Understanding, Generating, and Adapting User Interfaces","url":"https://www.collaborative-ai.org/publications/zhang22_caugaui/","source_class":"PRIMARY_RESEARCH","publication_date":"2022","accessed_at":"2026-08-03","claims_supported":["A 15-participant study predicted next actions and higher-level intents from interaction histories.","Action sequences achieved reported next-action accuracy of 66% and intent accuracy of 96% in a bounded text-formatting task.","The result supports limited technical plausibility for task-scoped next-action prediction but is too small and narrow to establish performance in enterprise workflows."]},{"source_id":"S7","title":"Privacy Principles","publisher":"World Wide Web Consortium","url":"https://www.w3.org/TR/privacy-principles/","source_class":"STANDARD","publication_date":"2025-05-15","accessed_at":"2026-08-03","claims_supported":["The W3C Statement requires data minimization, purpose limitation, transparency, meaningful consent, and easy withdrawal or objection.","The statement specifically recognizes that JavaScript can record interaction patterns and that behavioral information can be sensitive.","The candidate therefore requires consent-scoped capture, protected exclusions, withdrawal handling, retention limits, and privacy review; the statement does not establish jurisdiction-specific legal compliance."]},{"source_id":"S8","title":"Web Developers and Digital Designers","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/ooh/computer-and-information-technology/web-developers.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2025","accessed_at":"2026-08-03","claims_supported":["The May 2024 median wage was $98,090 for web and digital-interface designers and $90,930 for web developers.","These wages provide a public labor-cost anchor for broad 2026 resource-equivalent estimates.","The source does not cover UX-research, privacy-review, accessibility-review, cloud, or commercial replay-platform costs."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible across independent current products: Hotjar supplies filters intended to narrow replay review, while LogRocket explicitly states that teams may watch hundreds of sessions and spend hours or days locating useful moments. AXNav independently reports tedious, overwhelming manual accessibility testing. These sources support review-attention scarcity and the value of contextual replay, although no independent prevalence estimate or realized-impact measurement was found.","source_ids":["S1","S4","S5"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"Product, engineering, support, UX, and accessibility-testing teams are identifiable adopters. LogRocket names Afp Capital as an early adopter and records an expressed interest in behavioral summaries; AXNav studied 10 professional accessibility QA staff who found structured replay useful. Evidence is nevertheless tied to adjacent products and studies, not a commitment to adopt this candidate's synchronized residual architecture.","source_ids":["S4","S5"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"LogRocket Galileo Highlights and Issues","similarity":"Watches session streams, surfaces impactful struggle and behavioral patterns, creates summaries, and links reviewers to relevant moments instead of requiring indiscriminate replay.","remaining_difference":"The opened source does not describe a declared next-action model acting as a synchronized encoder/decoder, reconstructible structured residuals, random raw-session audits, version handshakes, or automatic full-state fallback.","source_ids":["S4"]},{"name":"Hotjar filtered recordings","similarity":"Prioritizes sessions using rage clicks, U-turns, errors, paths, drop-offs, and other event filters so reviewers can focus on relevant behavior.","remaining_difference":"It is rule-based filtering of retained recordings, not prediction-relative residual representation with model revision and reconstruction guarantees.","source_ids":["S1"]},{"name":"Fullstory frustration signals","similarity":"Surfaces predefined anomalous interaction moments and permits scoped searches and suppression of known false-positive elements.","remaining_difference":"Its heuristics can mistake intended behavior for frustration and do not establish task-flow prediction, raw-audit sampling, or model-state synchronization.","source_ids":["S3"]},{"name":"Microsoft Clarity recordings, filters, and AI summaries","similarity":"Captures actions and HTML, reconstructs sessions, exposes event timelines, filters and segments, and provides AI summaries.","remaining_difference":"It retains a conventional full-session representation and the source does not describe suppressing expected task stretches as reconstructible residuals against a shared model.","source_ids":["S2"]},{"name":"AXNav chaptered accessibility-test replay","similarity":"Creates navigable recordings, flags potential issues, and directs professional reviewers' attention while preserving replay context.","remaining_difference":"It replays specified automated tests rather than modeling natural users' next actions or compressing consented research sessions into residual episodes.","source_ids":["S5"]}],"distinctive_claim_remaining":"For one stable task and application version, a versioned next-action/state model plus structured residual episodes, periodic anchors, independent random and accessibility-stratified full-session audits, and automatic fallback can preserve consequential and protected interaction evidence while yielding more confirmed breakdown evidence per reviewer-hour than full replay, rule-based frustration filtering, and current AI-summary prioritization after all modeling and audit labor is counted.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Existing products already capture event streams, reconstruct sessions, filter events, and summarize recordings. AXNav demonstrates navigable, annotated replay for professional reviewers, and a small task-specific study demonstrates that next-action prediction from action histories is technically possible. However, reported next-action accuracy was only 66% in a 15-person text-formatting study, and no source validates synchronized residual reconstruction, accessibility-stratified miss rates, fallback reliability, or net reviewer savings for enterprise task sessions. W3C principles make deployment conditional on meaningful consent, withdrawal, minimization, purpose limitation, and sensitive-data controls; jurisdiction-specific legal authority remains unverified.","source_ids":["S2","S3","S5","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Review-time scarcity is explicit and affects issue discovery, but neither prevalence nor downstream product impact is independently quantified.","source_ids":["S1","S4","S5"]},"stakeholder_pull":{"score":4,"rationale":"Multiple commercial tools serve the workflow, an early adopter is named, and professional accessibility testers expressed usefulness; pull for the exact residual architecture is untested.","source_ids":["S4","S5"]},"incremental_advantage":{"score":2,"rationale":"AI summaries, issue surfacing, filters, frustration signals, and navigable annotations already address much of the claimed attention benefit. The remaining advantage depends on untested reconstruction, audit, and fallback properties.","source_ids":["S1","S2","S3","S4","S5"]},"distinctiveness_plausibility":{"score":3,"rationale":"No opened source combined synchronized task-flow prediction, reconstructible residual episodes, independent raw audits, and automatic decompression, but the functional objective substantially collides with established products. World novelty is unmeasured.","source_ids":["S1","S2","S3","S4","S5"]},"technical_implementability":{"score":3,"rationale":"Capture, reconstruction, prioritization, and next-action prediction components exist, but prediction accuracy, synchronization, protected-event recall, and net savings are not demonstrated together.","source_ids":["S2","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"A UX lead could authorize a non-production shadow study, but production use requires user consent and privacy, accessibility, data-retention, and model-owner authority not established by the external evidence.","source_ids":["S4","S5","S7"]},"evidence_readiness":{"score":2,"rationale":"A bounded protocol is available, but decisive evidence requires proprietary consented sessions, blinded human classifications, injected faults, and live or offline system testing unavailable from web research.","source_ids":["S3","S5","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Independent full-session audits, protected accessibility strata, and fallback directly address acknowledged heuristic false positives and privacy-sensitive interaction capture, but their effectiveness remains untested.","source_ids":["S3","S5","S7"]},"scalability":{"score":3,"rationale":"Existing replay and AI-summary systems operate at session scale, but the candidate requires a maintained model and audit regime for each task and application version, creating uncertain scaling costs.","source_ids":["S2","S4","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Pre-registration, one workflow model, offline instrumentation, 60-120 previously consented or scripted sessions, blinded review, fault injection, and analysis without replacing full replay.","confidence":"LOW","assumptions":["Uses an existing replay/event-capture stack and one non-production workflow.","Requires roughly 0.1-0.25 engineer-year plus part-time UX, privacy, and accessibility review.","BLS 2024 wage medians are adjusted approximately to 2026 and loaded for benefits and overhead.","No participant recruitment, vendor licensing, or legal-remediation expense is included beyond a modest contingency."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Build the task-flow predictor, structured residual schema, replay buffer, reviewer interface, model/version controls, audit sampler, and study-specific governance for one workflow.","confidence":"LOW","assumptions":["Approximately 3-10 loaded professional-months across engineering, UX research, data science, privacy, and accessibility.","Existing event capture, identity, storage, and consent infrastructure can be reused.","Commercial platform fees and unusually complex enterprise procurement are excluded because public verified prices were not found."],"source_ids":["S2","S7","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production hardening for a bounded product area, including security and privacy review, consent and deletion workflows, accessibility validation, monitoring, fallback drills, integrations, support, and launch evaluation.","confidence":"LOW","assumptions":["Approximately 1.5-5 loaded professional-years across several roles.","Raw fallback storage, monitoring, incident response, and independent audits are retained.","The band excludes organization-wide deployment and jurisdiction-specific litigation or regulatory remediation."],"source_ids":["S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Operate and maintain one bounded workflow: model refreshes, application-version updates, sampling and adjudication, privacy/accessibility review, storage and compute, monitoring, and incident drills.","confidence":"LOW","assumptions":["Roughly 0.3-1.25 loaded professional-years plus moderate infrastructure for one workflow.","Costs rise into higher bands if extended across many workflows or if full-session fallback is frequent.","No verified public vendor price or workload benchmark was found."],"source_ids":["S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent product and research publishers explicitly describe excessive replay/manual-review effort and mechanisms for narrowing attention.","source_ids":["S1","S4","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"LogRocket identifies Afp Capital as an early adopter of automated behavioral summaries and quotes its IT-services manager's intended uses; AXNav studied professional accessibility QA staff who expressed workflow usefulness.","source_ids":["S4","S5"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim can be tested as noninferior protected-event recall and reconstruction fidelity plus superior confirmed-breakdown yield per total reviewer-hour against full replay, heuristic filtering, and AI-summary prioritization.","source_ids":["S1","S2","S3","S4"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single-task, single-version, offline shadow comparison with untouched holdout sessions, fixed review budgets, injected failures, raw-session adjudication, and explicit falsifiers is bounded.","source_ids":["S2","S3","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"A shadow study appears governable, but meaningful consent, withdrawal, sensitive-field exclusion, accessibility coverage, retention, and jurisdiction-specific legal authority must be approved before using real sessions. W3C principles are authoritative guidance, not proof of legal compliance.","source_ids":["S3","S5","S7"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"Labor scope can be anchored to BLS wages, but no verified vendor pricing, infrastructure-volume benchmark, UX-research wage benchmark, or privacy-review cost was found; all bands therefore have low confidence.","source_ids":["S8"]}},"next_evidence_step":"Pre-register an offline shadow study for one non-production enterprise task and one frozen application version using 60-120 previously consented research sessions or scripted equivalents. Retain every full trace as adjudication evidence. Fit only on the training partition and keep an untouched holdout. Have at least two reviewers, blinded to selection method, work equal time budgets under four conditions: random full replay with skip-inactivity, Hotjar/Fullstory-style heuristic filtering, an available AI-summary/issue-prioritization tool, and the candidate's residual episodes. Independently dual-review all holdout full sessions to establish a reference set of breakdowns, legitimate alternative routes, and protected accessibility interactions. Primary tests are confirmed breakdowns per total labor-hour and recall of the reference set; secondary tests are false-positive rate, accessibility-stratum recall, event-order reconstruction error, fallback frequency, reviewer agreement, and total modeling/audit cost. Inject a dropped event, missing heartbeat, version mismatch, sensitive-field boundary, alternate valid route, and accessibility-navigation sequence. Falsify the incremental claim if the candidate fails to preserve at least 95% of adjudicated consequential and protected events, exceeds one event of ordering error in more than 5% of sessions, falls back in more than 10% of ordinary sessions, systematically overselects an accessibility or legitimate-strategy stratum, or fails to improve confirmed breakdowns per total labor-hour by at least 25% over the strongest comparator.","blocking_evidence":["No candidate-specific comparison against full replay, heuristic filtering, or current AI-summary prioritization exists.","No independently adjudicated dataset establishes breakdown, alternative-route, or accessibility-event recall.","No evidence validates model/decoder synchronization, raw-audit sampling rates, or fallback behavior under event loss and version mismatch.","No production authority, jurisdiction-specific legal basis, consent language, retention schedule, or deletion mechanism has been externally verified.","No evidence shows that modeling, synchronization, audit, and review costs remain below the reviewer time saved.","No public cost evidence covers commercial replay licensing, cloud volume, privacy review, accessibility review, or UX-research labor."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes substantial collision with session-replay filtering, frustration detection, AI issue surfacing and summarization, structured replay, and next-action prediction. It did not measure world novelty, patentability, freedom to operate, market size, or realized impact, and it was not an exhaustive literature, product, standards, or patent search. The only remaining claim is the bounded contrastive performance and governance claim stated here.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure privacy- and accessibility-approved access to one frozen workflow and consented or scripted full-session dataset.","Pre-register reconstruction, protected-event recall, reviewer-yield, labor-cost, subgroup, and fallback thresholds before fitting the predictor.","Run blinded equal-budget comparisons against full replay, rule-based frustration filtering, and an available AI-summary/issue-prioritization system.","Independently adjudicate raw holdout sessions and report accessibility-mode and legitimate-alternative-route results separately.","Demonstrate consent withdrawal, sensitive-field exclusion, event-loss detection, version-mismatch rejection, and safe fallback or halt.","Replace low-confidence cost assumptions with observed engineering, review, audit, storage, compute, and licensing resource use."],"reason":"Web evidence verifies a material review-attention problem, credible adopters, implementable components, and substantial prior-art collision. The decisive incremental claim—safe reconstruction with protected-event recall and superior evidence yield after total costs—requires proprietary consented interaction traces, blinded human review, fault injection, and system testing; bounded web research cannot answer it."},"proposal_index":2}