{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"negative_space_design__rhetoric:P1:v0","cell_id":"negative_space_design__rhetoric","search_queries":["site:census.gov cognitive interviewing think aloud verbal probing questionnaire evaluation moderator neutral","site:gov.uk focus group moderator silence leading questions guidance","focus group moderation wait time silence leading questions research study","message testing focus group moderator bias leading questions official guidance","focus group moderator manual silence pause probing write individual response official PDF","cognitive interview guide silence wait time neutral probe moderator official","site:aapor.org qualitative research focus groups consent recording standards","site:cdc.gov message testing focus groups moderator guide communication testing","experimental study interviewer probes leading cognitive interview reactivity response primary research","cognitive interviewing reactivity think aloud probing affects responses study","focus group moderator influence participant responses empirical study moderator bias","silence social pressure interview participants empirical research pause response","User Interviews participant recruitment pricing focus groups 2026 official","Respondent research participant recruitment pricing official focus group","focus group transcription pricing per minute official 2026","focus group moderator consultant rates cost guide 2025","\"Probing in Cognitive Interviews can Promote Acquiescence\" DOI","\"Pause for effect\" 10-s interviewer wait time DOI full text","41953102 probing cognitive interviews authors journal 2025"],"sources":[{"source_id":"S1","title":"Message testing","publisher":"World Health Organization","url":"https://www.who.int/initiatives/epi-win/the-collective-service/message-testing","source_class":"OFFICIAL_GUIDANCE","publication_date":"Not stated","accessed_at":"2026-08-03","claims_supported":["Communication teams are advised to test messages with intended audiences for comprehension, strengths, weaknesses, relevance, and sensitive or confusing elements.","WHO identifies focus groups and in-depth interviews as feasible message-testing methods and identifies partner organizations as possible recruiters or testing partners.","The 5-5-5 guide demonstrates institutional demand for bounded, low-cost message testing."]},{"source_id":"S2","title":"Small-Scale Testing: Fast, easy, free ways to test health messages and materials","publisher":"U.S. Centers for Disease Control and Prevention and Agency for Toxic Substances and Disease Registry","url":"https://www.cdc.gov/nceh/clearwriting/docs/Small-Scale-Testing-Guide-508.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"Not stated","accessed_at":"2026-08-03","claims_supported":["CDC identifies public-health practitioners, federal employees, contractors, and grantees as users and authorizers of message-testing protocols.","CDC recommends formal testing protocols, data-management plans, consideration of safety and ethics, and audience tests including interviews, focus groups, paraphrase tests, and comparative tests.","Its paraphrase procedure directs interviewers not to correct or interrupt participants before recording interpretations.","Its sample focus-group guide requires consent, voluntary participation, permission to pass or stop, recording disclosure, and assurance that the material—not the participant—is being tested.","For U.S. federal work, identical questions asked of 10 or more people may trigger Paperwork Reduction Act review; the guide advises consultation with the relevant PRA or project officer."]},{"source_id":"S3","title":"How to Organise and Run Focus Groups","publisher":"UK Health and Safety Executive","url":"https://www.hse.gov.uk/stress/assets/docs/focusgroups.pdf","source_class":"OFFICIAL_GUIDANCE","publication_date":"Not stated","accessed_at":"2026-08-03","claims_supported":["Official focus-group guidance says un-cued questions should precede cued questions.","It directs facilitators to avoid leading questions containing an implied answer and to use open questions that do not imply an expected response.","It describes neutral probes as an established way to elicit more detail."]},{"source_id":"S4","title":"Using Focus Groups in Program Development and Evaluation","publisher":"University of Kentucky Cooperative Extension","url":"https://psd.ca.uky.edu/files/focus.pdf","source_class":"AUTHORITATIVE_SECONDARY","publication_date":"Not stated","accessed_at":"2026-08-03","claims_supported":["A five-second silent pause after a moderator question is an established focus-group technique for encouraging a response.","Neutral probes and avoidance of approving judgments are established moderation practices.","Recording, transcription, and coding of focus-group responses are standard implementation components."]},{"source_id":"S5","title":"Appendix A2: Questionnaire Testing and Evaluation Methods for Censuses and Surveys","publisher":"United States Census Bureau","url":"https://www.census.gov/about/policies/quality/standards/appendixa2.html","source_class":"STANDARD","publication_date":"Not stated","accessed_at":"2026-08-03","claims_supported":["Cognitive interviewing, think-aloud reporting, paraphrasing, probing, respondent debriefing, split-panel tests, and behavior coding are established pretesting practices.","Census guidance notes that focus-group interaction does not provide a good test of an individual's response process when alone and that dominant participants can restrict others' input.","The standard supports iterative small-sample pretesting and recognizes different strengths and weaknesses among methods."]},{"source_id":"S6","title":"Probing in Cognitive Interviews can Promote Acquiescence","publisher":"Methods, Data, Analyses","url":"https://pubmed.ncbi.nlm.nih.gov/41953102/","source_class":"PRIMARY_RESEARCH","publication_date":"2025-10-08","accessed_at":"2026-08-03","claims_supported":["In a randomized probing experiment, 67 analyzed interviews compared directive and non-directive cognitive probes.","Respondents receiving directive probes affirmed the suggested interpretation more than five times as often as non-directively probed respondents volunteered it.","The experiment directly supports the possibility that interviewer-supplied candidate interpretations can contaminate evidence about respondent-originated interpretations, although it did not study live policy warrants or moderator silence."]},{"source_id":"S7","title":"Pause for effect: A 10-s interviewer wait time gives children time to respond to open-ended prompts","publisher":"Journal of Experimental Child Psychology","url":"https://pubmed.ncbi.nlm.nih.gov/32127193/","source_class":"PRIMARY_RESEARCH","publication_date":"2020-02-29","accessed_at":"2026-08-03","claims_supported":["Analysis of 105 conversations found that interviewers could usually follow a ten-second wait rule.","Children sometimes needed more than five seconds, and more than 96% of pauses followed by event information ended within ten seconds.","The study supports technical feasibility of a bounded wait protocol but concerns children recalling events, not adults evaluating policy arguments; it does not validate twelve seconds for the proposed setting."]},{"source_id":"S8","title":"Research Participant Recruitment Pricing","publisher":"Respondent","url":"https://www.respondent.io/pricing","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"Current pricing; publication date not stated","accessed_at":"2026-08-03","claims_supported":["Displayed 2026 pricing lists pay-as-you-go recruitment at approximately $15 per completed B2C session and $30 per completed B2B session, with lower committed rates.","Participant incentives are separate and researcher-funded.","The service includes screening, recordings, scheduling, and synthesis functions, establishing that remote recruitment and recording infrastructure is commercially available; final pricing remains commitment- and audience-dependent."]}],"problem_evidence":{"support":"MODERATE","rationale":"The underlying contamination mechanism is visible and consequential: official guidance warns against leading or cued questions, Census guidance says group interaction does not reveal an individual's response process, and a primary experiment found large differences between directive and non-directive probing. However, no source measured how often live policy-message moderators fill the first silence, whether participants then reuse moderator language, or how often this changes real communication decisions. The candidate-specific prevalence and impact magnitude therefore remain unverified.","source_ids":["S3","S5","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"WHO and CDC identify public-information and public-health communication teams, contractors, grantees, research offices, and project officers as actors who conduct or authorize message testing, and explicitly express the need to test comprehension, interpretation, actionability, and discomfort. CDC supplies a moderator guide for testing PSAs. These are credible adopter and authorizer classes, but no named organization expressed demand for this particular twelve-second warrant protocol or committed to adopting it.","source_ids":["S1","S2"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"Five-second focus-group pause and neutral probe","similarity":"Directly uses moderator silence after a question, followed by neutral probes, to let participants produce additional material without moderator judgment.","remaining_difference":"It is a general elicitation technique rather than a claim-warrant diagnostic; it does not specify a thinking-time cue, private first-response capture, delayed recovery of an intended rationale, origin coding, or a twelve-second maximum.","source_ids":["S4"]},{"name":"Un-cued-before-cued focus-group questioning","similarity":"Withholds answer cues and leading content until participants have responded freely, matching the proposed ordering logic.","remaining_difference":"It does not reserve a timed silent interval or define a claim-specific warrant-origin outcome and falsifier.","source_ids":["S3"]},{"name":"Cognitive interviewing, paraphrase testing, and behavior coding","similarity":"Established methods capture participant interpretations, use think-aloud or paraphrasing, record interviewer-participant interactions, and compare or code behavior before instrument deployment.","remaining_difference":"These methods generally use active verbalization or probes and are not specifically configured to test whether an audience independently supplies a policy argument's suppressed warrant during a framed silence.","source_ids":["S2","S5"]},{"name":"Non-directive versus directive probing experiment","similarity":"Directly tests whether interviewer-provided candidate interpretations contaminate evidence about a respondent's own interpretation.","remaining_difference":"It compares probe wording in cognitive interviews rather than a live silence protocol, private response option, or delayed disclosure of a policy rationale.","source_ids":["S6"]},{"name":"Ten-second interviewer wait guideline","similarity":"Implements and evaluates a bounded interval during which interviewers refrain from adding another prompt.","remaining_difference":"The population was children recalling events; it did not test adult policy reasoning, a visible silence cue, perceived pressure, warrant provenance, or the proposed twelve-second limit.","source_ids":["S7"]}],"distinctive_claim_remaining":"Not a claim of world novelty: compared with ordinary semi-structured moderation and a private written-before-discussion comparator, a claim-specific package combining a neutral warrant question, explicitly framed silence of at most twelve seconds, participant-controlled response modes, capture before explanation, and blinded provenance coding will increase the proportion of warrants whose participant-versus-moderator origin can be classified unambiguously without increasing misunderstanding, perceived pressure, or inability to pass.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The component techniques—scripts, pauses, neutral questions, recording, transcription, behavior coding, private or paraphrased responses, consent, passes, and comparative testing—are routine and require no new technology. Ten-second waiting has been operationally feasible in another interview population, and commercial remote research infrastructure is available. The main unresolved issues are measurement validity, moderator adherence, cultural and power-related interpretations of silence, accessibility, data protection for recordings, and whether twelve seconds is appropriate for adults evaluating policy arguments. Federal teams must also check PRA and organizational review requirements; larger or sensitive studies may require research-office or ethics review.","source_ids":["S2","S4","S5","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Preventing moderator-seeded warrants could avert false-positive message decisions, and primary evidence shows directive probes can substantially alter reported interpretations. Candidate-specific prevalence, downstream decision effects, and realized impact are unknown.","source_ids":["S5","S6"]},"stakeholder_pull":{"score":3,"rationale":"WHO and CDC visibly need message-testing methods and identify relevant organizational users, but there is no expressed demand or commitment for this exact protocol.","source_ids":["S1","S2"]},"incremental_advantage":{"score":2,"rationale":"Silence, nonleading questions, response capture, paraphrasing, and interaction coding are established. The remaining advantage is a more explicit provenance-measurement bundle, not a new core method, and it has not outperformed existing alternatives.","source_ids":["S2","S3","S4","S5"]},"distinctiveness_plausibility":{"score":2,"rationale":"The claim-specific combination is narrower than the located practices, but its central causal move—waiting silently and withholding cues before probing—is established practice. World novelty was not measured.","source_ids":["S3","S4","S5"]},"technical_implementability":{"score":5,"rationale":"The protocol uses ordinary facilitation, timers, response cards or forms, recordings, timestamps, and coding. No new technical capability is required.","source_ids":["S2","S4","S7","S8"]},"adoption_authority_feasibility":{"score":4,"rationale":"Communication research leads ordinarily control moderator scripts, while participants retain authority to pass or stop. Federal PRA, research-office, recording-consent, and sensitive-topic reviews must be handled where applicable.","source_ids":["S1","S2"]},"evidence_readiness":{"score":4,"rationale":"A small comparative feasibility test can be run immediately with matched claims, recordings, blinded coders, and predeclared discomfort and attribution outcomes. The coding instrument itself still needs reliability testing.","source_ids":["S2","S5","S6","S7"]},"safety_net_benefit":{"score":3,"rationale":"The protocol could protect against interviewer-created agreement and preserves pass, stop, and clarification options. Silence may itself create pressure, so the safety benefit depends on framing, a short maximum, accessibility choices, and immediate termination on request or distress.","source_ids":["S2","S6","S7"]},"scalability":{"score":4,"rationale":"A short script and coding template can diffuse across message-testing teams, and recruitment and recording infrastructure already exists. Moderator training, quality assurance, multilingual validation, and analysis labor limit frictionless scaling.","source_ids":["S1","S2","S4","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"UNDER_10K","scope":"An eight-adult, two-session remote feasibility and measurement check using three matched non-sensitive claims, three counterbalanced conditions, one trained moderator, recordings, two blinded coders, and an anonymous pressure/clarity survey.","confidence":"MODERATE","assumptions":["Existing videoconferencing and secure storage are available.","Eight participants receive assumed incentives of $75-$150 each; this incentive assumption is not directly verified by S8.","Approximately 32-60 loaded staff hours cover protocol preparation, moderation, coding, and a short report.","External legal review, specialist recruitment, travel, and effect-size estimation are excluded.","If conducted by a U.S. federal agency, the design remains below ten people or obtains required PRA guidance."],"source_ids":["S2","S4","S8"]},"initial_deployment_startup":{"band_2026_usd":"10K_TO_50K","scope":"Convert the method into an organizational protocol: research and privacy review, moderator training, accessibility and multilingual variants, coding manual, secure recording workflow, preregistration template, and one target-audience pilot.","confidence":"MODERATE","assumptions":["Approximately 120-300 loaded staff or consultant hours.","Uses existing meeting and storage infrastructure.","Includes limited target-audience recruitment and incentives but no custom software.","Sensitive, compulsory, crisis, trauma, or clinical contexts remain excluded."],"source_ids":["S2","S4","S8"]},"operational_launch":{"band_2026_usd":"10K_TO_50K","scope":"Use the approved protocol in one consequential communication project with three to six sessions, approximately 24-48 participants, independent transcription or timestamping, double coding, quality assurance, and a decision memo.","confidence":"LOW","assumptions":["Participant recruitment and incentives vary materially by audience.","Recruitment platform fees exclude incentives.","Professional moderation or specialist recruitment may dominate cost.","Any required PRA, IRB, privacy, procurement, or translation work stays within this band; extensive review could exceed it."],"source_ids":["S1","S2","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain quarterly-to-monthly use across multiple message projects, including recruitment, incentives, moderator calibration, recordings, double coding, accessibility support, protocol updates, audit sampling, and synthesis.","confidence":"LOW","assumptions":["Approximately 4-12 projects per year with 3-6 sessions each.","Existing staff and enterprise research infrastructure are shared across projects.","No dedicated software development or national representative sampling is included.","Specialized professional audiences, many languages, in-person travel, or outsourced full-service research could move costs above the band."],"source_ids":["S2","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Official moderation guidance identifies leading and cued questions as threats, Census guidance identifies limits on observing individual response processes in groups, and a primary randomized probing study demonstrates interviewer-supplied interpretations can alter reported interpretations. Exact prevalence in policy-message sessions remains unknown but the causal problem is externally supported.","source_ids":["S3","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"WHO and CDC explicitly address communication practitioners conducting audience message tests; CDC also identifies federal project officers, PRA contacts, and organizational research offices as authorizers where applicable.","source_ids":["S1","S2"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim compares a framed, bounded, pre-explanation warrant-capture package against ordinary moderation and private written capture on provenance classification, comprehension, pressure, and passability. It is testable even though the core pause practice is established.","source_ids":["S2","S4","S5","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"An eight-person, two-session, counterbalanced three-condition feasibility check can produce timestamps, blinded provenance labels, coder agreement, comprehension results, and participant-reported pressure without claiming an effect estimate.","source_ids":["S2","S5","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"A non-sensitive adult mock can proceed with informed recording consent, voluntary participation, explicit write/speak/pass options, a short maximum, and immediate termination for clarification requests, distress, or accessibility barriers. Federal teams must stay below the PRA threshold described by CDC or obtain authorization; consequential or sensitive settings remain excluded.","source_ids":["S2"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The scopes specify participant counts, sessions, staffing, analysis, exclusions, and escalation assumptions. Current first-party recruitment pricing anchors a small part of cost, while broad bands and low-to-moderate confidence appropriately reflect unverified incentives, labor, review, and specialist-recruitment costs.","source_ids":["S2","S8"]}},"next_evidence_step":"Run two recorded mock groups of four consenting adults using three matched, non-sensitive policy claims. Within each group, counterbalance claim-to-condition assignment across: (A) ordinary semi-structured moderation with no prescribed wait, (B) the Protected Warrant Silence package with the neutral question, visible cue, up-to-12-second interval, and write/speak/pass options, and (C) private written response captured before discussion. Preserve timestamps and exact first responses. Two coders blind to condition should independently classify whether a warrant appeared before moderator explanation, its apparent source, comprehension of the claim, and any moderator deviation. Collect anonymous 1-5 pressure and clarity ratings plus preferred response mode. Treat this only as feasibility and measurement testing. Falsify or halt the approach if coder agreement on origin is below 0.60, condition B does not yield more unambiguously attributable warrants than A, condition C is at least as clear with lower pressure, comprehension worsens, any participant cannot readily pass or request clarification, or the framed silence produces distress or a median pressure increase of at least one scale point.","blocking_evidence":["No recording audit establishes how frequently moderators prematurely supply warrants in real policy-message testing or how often doing so changes decisions.","No adult live-argument comparison shows that the proposed package improves warrant-origin attribution over ordinary nonleading moderation or private written capture.","The twelve-second maximum and thinking-time cue have not been validated across adult populations, languages, cultures, disabilities, or power relationships.","The proposed provenance-coding scheme has no demonstrated inter-rater reliability or construct validity.","No named communication organization has committed to adopting the specific protocol.","Downstream effects on message quality, public understanding, behavior, or harm remain unmeasured."],"research_disposition":"KNOWN_PRACTICE_DIFFUSION","world_novelty_boundary":"The search establishes only that the central practices—moderator pauses, nonleading and un-cued questions, pre-explanation interpretation capture, cognitive interviewing, and interaction coding—are already established. It does not establish exhaustive world novelty or non-novelty of the exact package. Patentability, freedom to operate, market size, realized impact, and worldwide practice outside the opened sources remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_DIFFUSION","repairable":false,"material_progress_observed":false,"progress_targets":["Reframe the proposal as a standardized implementation and measurement bundle built from established moderation and cognitive-interview practices, not as a novel silence technique.","Before consequential adoption, complete the eight-person three-condition feasibility test and demonstrate origin-coding agreement of at least 0.60 with no one-point median pressure penalty.","Predeclare that equal-or-better performance by private written capture, failure to beat ordinary moderation on unambiguous provenance, reduced comprehension, or any inability to pass falsifies the incremental advantage.","For any U.S. federal use, document PRA determination and organizational research/privacy approval; retain consent, recording, accessibility, and immediate-stop safeguards.","If diffusion proceeds, audit moderator adherence and participant interpretations of silence across languages and power contexts before scaling."],"reason":"The core intervention substantially overlaps established practice: focus-group manuals already prescribe silent pauses, official guidance requires un-cued and nonleading questions, and established cognitive-interview methods capture and code respondent interpretations. The claim-specific framing, delayed rationale, and provenance coding form a useful implementation bundle, but not enough externally supported distinctiveness remains for an innovation pipeline. Its incremental performance can only be resolved through live comparative testing, while the practical elements are suitable for cautious diffusion as known practice."},"proposal_index":1}