{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"catalytic_pathway_enablement__sociology_anthropology:P3:v0","cell_id":"catalytic_pathway_enablement__sociology_anthropology","search_queries":["site:dedoose.com training center test codes excerpts agreement official","site:help-nv.qsrinternational.com coder comparison query official NVivo kappa","site:atlasti.com intercoder agreement official documentation","team based qualitative coding calibration codebook intercoder agreement methodological study","site:atlasti.com/guides inter-coder agreement qualitative coding","site:dedoose.com pricing 2026 monthly official","site:ukdataservice.ac.uk consent qualitative data reuse research ethics official","site:nih.gov qualitative research team coding intercoder calibration codebook","qualitative multi-site study coding consistency calibration team analysts NIH implementation science PDF","site:obssr.od.nih.gov qualitative methods coding team researchers guide","site:hhs.gov/ohrp regulations identifiable private information secondary research qualitative data excerpts","site:ukdataservice.ac.uk qualitative data consent reuse sensitive data research teams","\"Embedding Big Qual and Team Science\" cross-site qualitative coding consistency","\"Preparing for analysis\" multisite qualitative research coding consistency expert coder","\"Intercoder Reliability in Qualitative Research\" O'Connor Joffe full text","qualitative coding calibration team methods expert coder transfer novel excerpts"],"sources":[{"source_id":"S1","title":"Codebook Development for Team-Based Qualitative Analysis","publisher":"Cultural Anthropology Methods / SAGE","url":"https://qualquant.org/wp-content/uploads/text/MacQueen%20et%20al%201998.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"1998","accessed_at":"2026-08-03","claims_supported":["CDC-based teams used standardized codebooks because multiple, geographically dispersed coders had to analyze large volumes of text.","Coders applied codes to common sample text, compared segmentation and code application, discussed inconsistencies, revised ambiguous definitions, and repeated agreement checks.","Periodic checking and codebook stewardship are established practices rather than novel elements of the candidate."]},{"source_id":"S2","title":"Structuring a Team-Based Approach to Coding Qualitative Data","publisher":"International Journal of Qualitative Methods / SAGE","url":"https://journals.sagepub.com/doi/10.1177/1609406920968700","source_class":"PRIMARY_RESEARCH","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["The authors observed coding-team confusion, mistakes, missing quality review, and unclear authority over coding-scheme changes.","Their implemented workflow used a lead analyst, external qualitative methodologist, coders, senior reviewers, training, arbitration, and meetings roughly every ten days during six weeks of coding.","Detailed codebooks and coded reference transcripts improved shared understanding, while the codebook remained a living document requiring revision."]},{"source_id":"S3","title":"Preparing for Analysis: A Practical Guide for a Critical Step for Procedural Rigor in Large-Scale Multisite Qualitative Research Studies","publisher":"Quality & Quantity / Springer Nature; archived by PubMed Central","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10614080/","source_class":"PRIMARY_RESEARCH","publication_date":"2017-02-28","accessed_at":"2026-08-03","claims_supported":["A NIDA-funded cooperative spanning nine research centers and about 700 interviews required coordinated codebook development, training, data preparation, and intercoder consistency procedures.","Codebook development involved lengthy rounds of discussion; training was ongoing, iterative, and repetitive.","Multisite coders calibrated on shared transcripts, discussed discrepancies, repeated exercises, escalated unresolved issues to workgroups or team leaders, and in one approach used an 80-percent agreement threshold.","Investigators, project managers, protocol-specific qualitative workgroups, team leaders, and qualitative experts are identifiable adopters and authorizers for this type of workflow."]},{"source_id":"S4","title":"Developing Shared Ways of Seeing Data: The Perils and Possibilities of Achieving Intercoder Agreement","publisher":"International Journal of Qualitative Methods / SAGE","url":"https://journals.sagepub.com/doi/pdf/10.1177/16094069231160973","source_class":"PRIMARY_RESEARCH","publication_date":"2023","accessed_at":"2026-08-03","claims_supported":["Analyst subjectivities, cultures, histories, expertise, and power relations materially affect team interpretation and consensus.","Disagreements can expose ambiguity or multiple reasonable interpretations rather than coder error.","Race, gender, class, institutional position, and other power differentials can constrain the ability to disagree, making apparently high consensus an unsafe proxy for truth.","The authors developed a structured alignment method while seeking to preserve high-inference interpretation, demonstrating both feasibility and the need for contextual safeguards."]},{"source_id":"S5","title":"Beyond a Coefficient: An Interactive Process for Achieving Inter-Rater Consistency in Qualitative Coding","publisher":"Qualitative Research / SAGE; ERIC full-text deposit","url":"https://files.eric.ed.gov/fulltext/ED643124.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2022","accessed_at":"2026-08-03","claims_supported":["Four coders calibrated by blindly coding shared observations and interviews, comparing code applications, discussing reasoning, expanding definitions with examples and non-examples, and escalating persistent disputes to co-principal investigators.","The workflow included a Not Sure code with explanatory memos, twice-weekly resolution meetings, midstream recalibration, and checks of within-coder consistency.","A traditional expert-master test was judged poorly suited to layered, context-dependent codes, directly supporting the candidate's warning against treating one interpretation as ground truth.","The reported process lasted about five months and involved 47 meetings, evidencing substantial recurring labor while also showing that most proposed workflow components are established practice."]},{"source_id":"S6","title":"Testing Center (IRR Using Cohen's Kappa)","publisher":"Dedoose Learning Center","url":"https://helpdesk.dedoose.com/hc/en-us/articles/12905437600269-Testing-Center-IRR-using-Cohen-s-Kappa","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated","accessed_at":"2026-08-03","claims_supported":["Dedoose already offers reusable tests in which administrators select codes and previously coded excerpts and trainees apply codes to those excerpts.","The product reports pooled and code-specific Cohen's kappa and excerpt-level agreement or disagreement against an expert coder.","This is a close product collision with the proposed anchor bank, bounded analyst packet, discrepancy profile, and readiness-calibration function.","Dedoose cautions that test validity depends on sufficient excerpts and recommends limiting tests to essential or frequent codes."]},{"source_id":"S7","title":"Measuring Inter-Coder Agreement — ATLAS.ti 26 Windows User Manual","publisher":"ATLAS.ti Scientific Software Development GmbH","url":"https://manuals.atlasti.com/Win/en/manual/ICA/ICAMeasuring.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026","accessed_at":"2026-08-03","claims_supported":["ATLAS.ti already measures coder agreement and disagreement over text, audio, and video.","Its implementation handles disagreement in both code application and selected segment boundaries and supports predefined quotations where coders only apply codes.","Existing product capability makes basic automated discrepancy calculation technically straightforward but does not itself establish contextual validity, readiness dispositions, specialist triage, or regeneration."]},{"source_id":"S8","title":"Coded Private Information or Biospecimens Used in Research, Guidance (2018)","publisher":"U.S. Department of Health and Human Services, Office for Human Research Protections","url":"https://www.hhs.gov/ohrp/coded-private-information-or-biospecimens-used-research.html","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2018","accessed_at":"2026-08-03","claims_supported":["Using, studying, or analyzing identifiable private information can constitute human-subjects research, and coding identifiers does not automatically remove that status.","Secondary research involving coded information may fall outside the human-subject definition only under conditions preventing investigators from readily ascertaining identities; otherwise exemption or IRB review may be required.","Reusing participant excerpts for calibration therefore requires a study-specific institutional determination and appropriate data-custodian authority; synthetic vignettes materially reduce but do not eliminate governance requirements for research involving analyst participants."]}],"problem_evidence":{"support":"STRONG","rationale":"Multiple independent applied studies document recurring coder confusion, lengthy and repetitive training, shared-transcript calibration, discrepancy adjudication, codebook revision, and continued drift checks. One detailed precedent required five months and 47 meetings, while a nine-center study described lengthy codebook development and repeated multisite calibration. These sources establish visible burden and methodological stakes, although they do not estimate prevalence across all comparative ethnography projects or quantify the candidate's prospective effect size.","source_ids":["S1","S2","S3","S5"]},"stakeholder_evidence":{"support":"STRONG","rationale":"Actual multisite studies identify principal investigators, project managers, qualitative workgroups, lead analysts, external methodologists, senior reviewers, and site coders as users and decision-makers. They expressly sought consistency, procedural rigor, clearer responsibility, and scalable handling of large qualitative datasets. No named organization has committed to adopt this particular circuit, so evidence supports credible stakeholder pull rather than procurement or adoption commitment.","source_ids":["S2","S3","S4"]},"prior_art":{"proximity":"ESTABLISHED_PRACTICE","closest_analogues":[{"name":"Dedoose Testing Center","similarity":"Very close product analogue: reusable expert-coded excerpt tests, trainee coding, code-specific discrepancy results, and an explicit purpose of building and maintaining inter-rater reliability.","remaining_difference":"It treats an expert coding as the comparison reference and does not document the candidate's contextual-plurality rules, local-expert review, scoped ready/revise/escalate disposition, independent transfer packet, workload regeneration, anchor quarantine, or nonpunitive governance.","source_ids":["S6"]},{"name":"Hemmler et al. interactive calibration and recalibration process","similarity":"Very close practice analogue: blind shared-case coding, reasoning discussion, examples and non-examples, escalation to co-PIs, uncertainty memos, repeated calibration, drift checks, and protection of contextual subjectivity.","remaining_difference":"It is meeting-intensive and does not test whether automated boundary-class profiling plus selective specialist review reduces total expert labor while maintaining performance on a distinct transfer packet.","source_ids":["S5"]},{"name":"MacQueen et al. team-based codebook calibration","similarity":"Foundational analogue using sample-text comparison, structured inclusion and exclusion criteria, discrepancy review, team-leader adjudication, codebook revision, and periodic agreement checks across dispersed coders.","remaining_difference":"It does not package these elements as an auditable readiness service with explicit capacity, regeneration, deactivation, contextual-plurality, and downstream-transfer metrics.","source_ids":["S1"]},{"name":"ATLAS.ti inter-coder agreement tooling","similarity":"Adjacent product capability for computing code and unit-boundary agreement or disagreement, including use of predefined quotations.","remaining_difference":"It measures agreement but does not provide the proposed governance, specialist triage, calibration-readiness disposition, regeneration cycle, or safeguards against false consensus.","source_ids":["S7"]}],"distinctive_claim_remaining":"Relative to ordinary team calibration and existing excerpt-test software, an integrated circuit that classifies boundary discrepancies, sends only stable operationalized boundary classes to a capped specialist lane, routes contextual or theoretical disagreements out of the circuit, and regenerates its anchors can reduce total senior and local-expert time and elapsed readiness time across repeated analyst-codebook pairings while remaining noninferior on a distinct transfer packet and preserving legitimate alternative readings. This is contrastive and falsifiable, but it is an untested system-level performance claim rather than a novel component claim.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"Every basic technical and workflow component has a precedent: versioned codebooks and examples, shared-case tests, rule-based agreement calculations, uncertainty memos, specialist escalation, repeated calibration, and governance by study leads. A synthetic-only shadow probe is technically straightforward. Evidence does not establish a validated discrepancy taxonomy, reliable transfer prediction, cross-language fairness, anchor-memorization controls, specialist-capacity restoration, or integration with a real study's IRB, data-use, labor, and community-governance requirements.","source_ids":["S1","S2","S3","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"On a five-point scale, documented multisite calibration burden, prolonged meeting load, and risks to interpretive trustworthiness make the target consequential; realized time savings and downstream analytic impact remain unmeasured.","source_ids":["S2","S3","S5"]},"stakeholder_pull":{"score":4,"rationale":"Multisite investigators and qualitative-methods leads visibly need scalable consistency procedures, but no adopter has requested or committed to this exact circuit.","source_ids":["S2","S3","S4"]},"incremental_advantage":{"score":2,"rationale":"The proposed integration could redirect scarce review and add regeneration and transfer safeguards, but its principal functions substantially overlap established calibration workflows and Dedoose's Testing Center.","source_ids":["S1","S5","S6","S7"]},"distinctiveness_plausibility":{"score":2,"rationale":"Distinctiveness survives only at the integrated operating-model level—selective specialist triage, contextual-plurality preservation, regeneration, and transfer validation—not at the component or basic workflow level.","source_ids":["S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"Existing software already presents excerpt tests and computes disagreements, while published teams already execute the human review and escalation steps. Context-sensitive classification and secure integration require validation but no apparent technical breakthrough.","source_ids":["S5","S6","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"PIs, qualitative workgroups, methods leads, data custodians, and IRBs are identifiable authorities. Feasibility depends on obtaining their coordinated approval and preventing readiness results from becoming unauthorized employment or substantive-interpretation decisions.","source_ids":["S2","S3","S8"]},"evidence_readiness":{"score":3,"rationale":"A small synthetic shadow comparison is well specified and can use established methods, but no prototype, partner, baseline logs, discrepancy taxonomy, or transfer-validation data currently exist.","source_ids":["S3","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"Synthetic anchors, explicit escalation, blinded local-context review, nonpunitive rollback, and deactivation can reduce privacy and false-consensus risk relative to uncontrolled calibration, provided local governance is secured.","source_ids":["S4","S8"]},"scalability":{"score":3,"rationale":"Reusable digital packets and targeted review could scale across repeated analyst-codebook pairings, but language, site context, codebook churn, local-expert compensation, anchor exposure, and specialist saturation limit generalization.","source_ids":["S3","S4","S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Design and run a synthetic-only, preregistered shadow study with at most 12 consenting analysts or trainees, two matched calibration pathways, two blinded reviewers, a distinct transfer packet, secure data capture, and analysis of labor, time, performance, escalation, plurality, and disparity outcomes.","confidence":"MODERATE","assumptions":["Existing survey, qualitative-analysis, or lightweight scripting tools are used rather than custom production software.","Approximately 200-400 total analyst, methods, local-context, engineering, governance, and analysis hours are valued as 2026 resource equivalents.","No participant-derived excerpts are used and analysts are compensated or participating within an authorized training role.","Institutional overhead and extensive multilingual translation are excluded."],"source_ids":["S3","S5","S6"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Create a governed minimum viable circuit: 30-60 synthetic or specifically authorized anchors, versioned specifications, discrepancy taxonomy, secure test interface, audit logs, capacity view, transfer packet, role definitions, data-use controls, and validation documentation.","confidence":"LOW","assumptions":["One organization and one established codebook are in scope.","The profiler is deterministic and auditable rather than an unvalidated machine-learning classifier.","Methods, local-context, data-protection, software, and project-management labor are fully counted.","No enterprise procurement, major systems integration, or new source-data collection is required."],"source_ids":["S2","S5","S6","S7","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"Launch one bounded study-site implementation, including workflow integration, analyst and specialist training, support, governance review, security testing, live shadow operation, independent audit, and rollback readiness before any access-control use.","confidence":"LOW","assumptions":["Launch covers one comparative study and fewer than 30 analysts.","Participant-derived material is excluded unless separately authorized.","The ordinary supervised pathway remains available during launch.","Costs include paid local-context review and specialist recovery time."],"source_ids":["S2","S3","S4","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Maintain one study's circuit through part-time methods stewardship, paid local-context review, anchor refresh and rotation, codebook synchronization, software and security support, incident handling, transfer sampling, fairness audits, and specialist coverage.","confidence":"LOW","assumptions":["Approximately 0.25-1.0 full-time-equivalent combined stewardship and support capacity is required.","Major codebook redesigns and additional languages or sites would raise costs.","Annual volume remains below specialist-lane saturation.","The band is a resource-equivalent estimate, not a vendor quote; public sources did not establish direct 2026 labor rates or total-cost benchmarks."],"source_ids":["S2","S3","S5","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Applied multisite studies independently document repeated training, shared-case calibration, discrepancy adjudication, codebook revision, drift checks, confusion, mistakes, and substantial meeting burden.","source_ids":["S1","S2","S3","S5"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"Published workflows identify PIs, project managers, qualitative workgroups, lead analysts, external methodologists, senior reviewers, data custodians, and IRBs as credible adopters or authorizers, with expressed needs for consistency and procedural rigor.","source_ids":["S2","S3","S8"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim compares the integrated circuit against ordinary calibration and existing excerpt-test practice on total expert labor, elapsed time, distinct-packet transfer, escalation quality, and preservation of acceptable interpretive plurality.","source_ids":["S5","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A counterbalanced synthetic shadow probe with at most 12 analysts, fixed comparators, blinded review, a distinct transfer packet, and precommitted superiority, noninferiority, fairness, privacy, and plurality falsifiers is bounded and reversible.","source_ids":["S3","S4","S5","S6"]},"no_unresolved_safety_or_authority_stop":{"status":"UNCERTAIN","reason":"A synthetic-only probe avoids unauthorized reuse of participant excerpts, but the actual institution, consenting analyst population, IRB or research-ethics determination, employment-use prohibition, data custodian, and local-context authority have not been named or secured.","source_ids":["S4","S8"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The four bands have explicit scopes and labor assumptions and are consistent with published workflows involving multiple specialist roles, repeated meetings, months of calibration work, and existing software capabilities. Confidence remains low to moderate because no adopter-specific rates or vendor implementation quote were found.","source_ids":["S2","S3","S5","S6","S7"]}},"next_evidence_step":"With a named comparative-study partner and local ethics or IRB determination, preregister a counterbalanced shadow probe involving no more than 12 consenting analysts or authorized trainees. Use only purpose-built synthetic vignettes in two stable code families and stratify analysts by language and site preparation. Randomize order between (A) the existing ordinary calibration procedure and (B) the circuit's bounded packet, rule-based discrepancy profile, and capped specialist review. Hold analyst eligibility, codebook, readiness standard, preparation, and total material constant. Blinded methods and local-context reviewers should score boundary reasoning, appropriate escalation, memo quality, and the number and quality of legitimate alternative readings on a distinct transfer packet never exposed during calibration. Count all analyst, senior, local-expert, tool-administration, retry, and meeting minutes. The incremental claim passes only if the circuit reduces total senior-plus-local-expert minutes per valid disposition by at least 25% and shortens elapsed disposition time, while the lower confidence bound excludes more than a five-percentage-point loss in transfer performance and shows no reduction in reviewer-accepted interpretive plurality. Falsify or halt on failure to meet either threshold, increased false readiness or unnecessary exclusion, material disposition disparity by language or site preparation, strategic answer matching, culturally misleading anchors, privacy or authorization breach, or inability to restore the anchor bank and specialist lane for a second cycle. Passing authorizes only a separately approved live shadow pilot, not analyst employment decisions or unsupervised access to participant data.","blocking_evidence":["No field measurement shows that the integrated circuit reduces total rather than merely visible senior and local-expert labor.","No evidence shows that its readiness disposition predicts coding on an unexposed transfer packet or a live batch.","No discrepancy taxonomy has been validated to distinguish routine boundary errors from productive theoretical, linguistic, or site-specific disagreement.","No named adopter has committed staff, supplied baseline process logs, or secured institutional ethics, data-custodian, community-governance, and employment-use authority.","No evidence establishes regeneration durability under repeated anchor exposure, codebook churn, reviewer fatigue, or multilingual use.","No adopter-specific labor rates, software integration quote, local-context compensation plan, or annual case volume supports a higher-confidence cost estimate."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search establishes extensive prior practice and close product analogues but does not measure world novelty. Patentability, freedom to operate, market size, and realized impact are also unmeasured. The only remaining evaluated distinction is the candidate's integrated, governed operating model and its untested comparative performance claim; no conclusion is made that this integration is globally novel.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure a named multisite or comparative qualitative-research partner and written PI, methods-governance, ethics or IRB, data-custodian, and nonpunitive-use determinations.","Build a minimum synthetic anchor set and auditable discrepancy taxonomy covering at least two stable code families, with local-context reviewers able to quarantine misleading cases.","Preregister and execute the no-more-than-12-analyst counterbalanced shadow probe against ordinary calibration.","Demonstrate at least a 25% reduction in total senior-plus-local-expert minutes per valid disposition while remaining within a five-percentage-point noninferiority margin on an unexposed transfer packet.","Show no reduction in reviewer-accepted interpretive plurality and no material disparity, privacy breach, culturally misleading anchor, false-readiness increase, or unauthorized employment use.","Repeat the circuit for a second cycle and document anchor exposure, restoration work, reviewer recovery time, predictive retention, and full resource-equivalent cost."],"reason":"Bounded web research materially resolved the problem, credible adopter roles, feasibility, authority boundary, and prior-art questions. It also found that shared-case calibration, discrepancy review, repeated recalibration, specialist escalation, and software-based excerpt tests are established practice, leaving only an integrated incremental performance claim. Whether that bundle actually lowers total labor, transfers to unseen material, preserves contextual plurality, and regenerates across cycles requires proprietary workflow data and a live controlled shadow test; under the controller rule this requires STOP_EMPIRICAL_RESEARCH_NEEDED, with repairable set false."},"proposal_index":3}