{"schema_version":1,"research_id":"eoa_inverse_innovation_exp04_external_evaluation_20260802","source_assessment_id":"computability_boundary_mapping__ethnography_qualitative_methods:PROPOSAL_FIRST:v0","cell_id":"computability_boundary_mapping__ethnography_qualitative_methods","search_queries":["qualitative research saturation stopping data collection guidance official","systematic review saturation qualitative research inconsistent reporting empirical study","executable qualitative research workflow software automation coding saturation","workflow language termination verification model checking Turing complete workflow prior art","site:nih.gov qualitative research saturation guidance rigor transparency sample size","site:equator-network.org SRQR qualitative research reporting standards saturation","site:dedoose.com workflow automation coding qualitative saturation official","site:lumivero.com NVivo automation API script qualitative coding official","O'Brien Standards for Reporting Qualitative Research SRQR 21 items full text","van der Aalst verification of workflow nets soundness primary paper PDF","Temporal official docs deterministic workflow code restrictions non deterministic","BLS Occupational Employment Wage Statistics software developers computer research scientists May 2025","site:pubmed.ncbi.nlm.nih.gov \"Standards for reporting qualitative research\" O'Brien 2014","site:equator-network.org/reporting-guidelines/srqr","qualitative data analysis software programmable workflow user authored scripts saturation validator","\"qualitative\" \"workflow\" \"model checking\" saturation","QDAcity theoretical saturation measurement paper official","site:qdacity.com theoretical saturation feature how works","QDAcity automated theoretical saturation qualitative research algorithm paper","\"theoretical saturation\" software tool workflow qualitative"],"sources":[{"source_id":"S1","title":"Saturation in qualitative research: exploring its conceptualization and operationalization","publisher":"Quality & Quantity / Springer Nature","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5993836/","source_class":"PRIMARY_RESEARCH","publication_date":"2018","accessed_at":"2026-08-02","claims_supported":["Saturation is used to decide when qualitative data collection or analysis can stop.","Saturation has multiple meanings and is used inconsistently across methodologies.","A saturation judgment is predictive and should be aligned with the research question, theoretical position, and analytic framework rather than treated as a universal mechanical fact."]},{"source_id":"S2","title":"Sample sizes for saturation in qualitative research: A systematic review of empirical tests","publisher":"Social Science & Medicine / Elsevier","url":"https://pubmed.ncbi.nlm.nih.gov/34785096/","source_class":"PRIMARY_RESEARCH","publication_date":"2021-11-02","accessed_at":"2026-08-02","claims_supported":["The review found 23 empirical or modeling studies and substantial dependence of saturation results on population homogeneity, study objectives, and whether code or meaning saturation was assessed.","The authors identify researchers, journal reviewers, ethical review boards, and funding agencies as users of more transparent saturation justification.","Reported interview ranges are conditional empirical findings, not universal stopping guarantees."]},{"source_id":"S3","title":"A simple method to assess and report thematic saturation in qualitative research","publisher":"PLOS ONE","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7200005/","source_class":"PRIMARY_RESEARCH","publication_date":"2020-05-05","accessed_at":"2026-08-02","claims_supported":["The paper offers an operational comparator based on base size, run length, and a new-information threshold.","The method is intended to make saturation assessment more transparent during or after collection.","The authors expressly state that meeting the proposed thresholds does not guarantee that saturation has in fact been reached."]},{"source_id":"S4","title":"Standards for reporting qualitative research: a synthesis of recommendations","publisher":"Academic Medicine / Wolters Kluwer","url":"https://pubmed.ncbi.nlm.nih.gov/24979285/","source_class":"STANDARD","publication_date":"2014-09","accessed_at":"2026-08-02","claims_supported":["SRQR comprises 21 reporting items intended to improve transparency across qualitative research.","The identified users include authors, journal editors, reviewers, and readers.","SRQR is reporting guidance rather than an executable termination or premature-stop verifier."]},{"source_id":"S5","title":"QDA Software | QDAcity","publisher":"QDAcity","url":"https://qdacity.com/en/","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"","accessed_at":"2026-08-02","claims_supported":["QDAcity is an identifiable cloud qualitative-data-analysis platform with collaboration, permissions, and automated transcription.","The product states that it helps measure theoretical saturation and provides inter-coder metrics.","The public product page does not state that users submit arbitrary executable workflows or that QDAcity universally decides workflow termination or substantive saturation."]},{"source_id":"S6","title":"Q-Sat AI: Machine Learning-Based Decision Support for Data Saturation in Qualitative Studies","publisher":"arXiv","url":"https://arxiv.org/abs/2511.01935","source_class":"PRIMARY_RESEARCH","publication_date":"2025-11-02","accessed_at":"2026-08-02","claims_supported":["This preprint proposes computational decision support for qualitative saturation and a future web application.","It targets researchers, journal reviewers, and thesis advisors and frames saturation ambiguity as a methodological-rigor problem.","Its reported machine-learning prediction is adjacent prior art but is not a proof of workflow termination, a computability classification, or an exact saturation certificate."]},{"source_id":"S7","title":"Soundness of workflow nets: classification, decidability, and analysis","publisher":"Formal Aspects of Computing / Springer Nature","url":"https://link.springer.com/article/10.1007/s00165-010-0161-4","source_class":"PRIMARY_RESEARCH","publication_date":"2010-08-03","accessed_at":"2026-08-02","claims_supported":["Workflow nets are a standard formal abstraction used to check soundness, including absence of deadlocks and livelocks.","The paper proves multiple soundness notions decidable for ordinary workflow nets.","It also shows that more expressive workflow-net extensions make most of those notions undecidable, directly supporting explicit language boundaries rather than a universal verifier."]},{"source_id":"S8","title":"National employment and wage data by occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026","accessed_at":"2026-08-02","claims_supported":["The national mean wage was $148,100 for software developers, $153,930 for computer and information research scientists, and $111,490 for software quality-assurance analysts and testers.","These wage benchmarks support labor-based resource-equivalent estimates, but not fixed-price quotes or institution-specific overhead."]}],"problem_evidence":{"support":"WEAK","rationale":"The methodological core is visible and consequential: saturation affects stopping, is inconsistently defined, and can lead to premature or excessive collection. However, no searched source documented the candidate's specific operational premise—a deployed qualitative platform accepting arbitrary executable workflows while advertising a universal Boolean termination and premature-stop validator. Publicly visible tools instead offer metrics, decision support, reporting guidance, or ordinary analyst-controlled workflows. The exact stated software problem therefore remains hypothetical or proprietary.","source_ids":["S1","S2","S3","S5","S6"]},"stakeholder_evidence":{"support":"WEAK","rationale":"Researchers, editors, reviewers, ethics boards, funders, readers, and thesis advisors are externally identified as users of transparent saturation evidence. QDAcity is an identifiable potential product partner already offering saturation measurement. None of these sources expresses demand for computability-boundary records, UNKNOWN routing, or formal termination checking, and no adopter commitment or authorizer interview was found.","source_ids":["S2","S4","S5","S6"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"QDAcity theoretical-saturation metrics","similarity":"A live qualitative-research platform already claims to help measure theoretical saturation and supplies workflow-adjacent metrics.","remaining_difference":"The public description does not expose arbitrary workflow code, universal termination validation, formal scope boundaries, proof artifacts, or exact/bounded/unknown guarantee labels.","source_ids":["S5"]},{"name":"Guest, Namey, and Chen thematic-saturation method","similarity":"Operationalizes an iterative qualitative stopping assessment using predeclared parameters and transparent reporting.","remaining_difference":"It is an empirical heuristic framework whose authors deny that thresholds guarantee saturation; it does not verify executable workflow termination or premature-stop safety.","source_ids":["S3"]},{"name":"Q-Sat AI decision support","similarity":"Uses computation to standardize saturation-related decisions and targets researchers and reviewers.","remaining_difference":"It is predictive decision support, not a computability analysis or sound formal verifier, and the searched record is a preprint rather than deployment evidence.","source_ids":["S6"]},{"name":"Workflow-net soundness verification","similarity":"Formally separates decidable workflow classes from expressive extensions with undecidable soundness properties.","remaining_difference":"It addresses generic formal workflows and domain-independent control-flow anomalies, not semantic fidelity to qualitative sampling, coding, negative-case analysis, or saturation.","source_ids":["S7"]},{"name":"SRQR reporting standard","similarity":"Requires transparent, reviewable qualitative-method reporting and serves authors, reviewers, editors, and readers.","remaining_difference":"It is a reporting standard without executable semantics, formal proof obligations, or runtime routing.","source_ids":["S4"]}],"distinctive_claim_remaining":"For a platform that demonstrably accepts executable qualitative-research workflows, an enforceable finite-state fragment plus a router can prevent every bounded, heuristic, timed-out, abstraction-based, or human-reviewed result from being emitted as universal exact certification, while returning EXACT, BOUNDED, UNKNOWN, or OUT_OF_SCOPE according to machine-checkable scope. This is falsified by any in-scope workflow receiving an incorrect exact label, any out-of-fragment workflow receiving exact certification, any lost UNKNOWN state, or evidence that the formal stopping property fails to represent the intended qualitative evidentiary condition.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Formal workflow research supports decidable restricted classes and shows that added expressiveness can cross into undecidability. Existing qualitative software demonstrates that cloud workflows and saturation metrics can be implemented. A synthetic finite-state DSL, enumerator, label router, and audit record are technically plausible. What is not established is semantic fidelity between formal stopping predicates and qualitative saturation, a reviewed halting reduction for the proposed language, usable guarantee labels, integration with a real platform, privacy controls for transcripts, or production performance.","source_ids":["S3","S5","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Premature or unjustified stopping can damage qualitative trustworthiness and omit perspectives, but prevalence and severity of the candidate's exact executable-validator failure are unknown.","source_ids":["S1","S2","S3"]},"stakeholder_pull":{"score":2,"rationale":"Multiple stakeholder classes want transparent saturation justification, yet no organization was found requesting the proposed computability-boundary intervention.","source_ids":["S2","S4","S5","S6"]},"incremental_advantage":{"score":3,"rationale":"Guarantee-aware labels and enforced scope would add a useful safeguard beyond saturation metrics, reporting standards, and timeout-based execution, conditional on a real executable-workflow use case.","source_ids":["S3","S4","S5","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"The cross-domain combination was not found, but each major element—saturation decision support, reporting transparency, and restricted workflow verification—has established adjacent prior art.","source_ids":["S3","S4","S5","S6","S7"]},"technical_implementability":{"score":4,"rationale":"The bounded synthetic prototype and finite fragment are credible; unrestricted exact validation is intentionally excluded. Production semantic modeling and state-space performance remain open.","source_ids":["S5","S7"]},"adoption_authority_feasibility":{"score":2,"rationale":"PIs and applicable ethics or governance bodies can authorize a research stopping decision, and a vendor could control platform labels, but no actual platform owner or governance body has agreed to the proposed allocation of authority.","source_ids":["S2","S4","S5"]},"evidence_readiness":{"score":2,"rationale":"The literature supports the general methodological issue and technical pattern, but candidate-specific prevalence, adopter demand, formal proof review, user interpretation, and production integration evidence are absent.","source_ids":["S1","S2","S5","S7"]},"safety_net_benefit":{"score":4,"rationale":"Preserving UNKNOWN and OUT_OF_SCOPE and forbidding software from declaring substantive saturation could reduce over-certification, provided users understand the labels and human review remains accountable.","source_ids":["S1","S3","S4"]},"scalability":{"score":3,"rationale":"A syntactically restricted DSL and automated routing can be reused across workflows, but coverage loss, state-space growth, and reviewer workload could limit scale.","source_ids":["S5","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Two-week synthetic exercise: formalize six workflows, implement a tiny finite-state checker and timeout baseline, draft the reduction, conduct independent proof-direction review, and write a guarantee record.","confidence":"MODERATE","assumptions":["Approximately 80 software-engineer hours, 40 qualitative-methodologist hours, 16 formal-methods review hours, and 16 project-management/documentation hours.","Loaded labor is assumed to be about 1.7-2.2 times cash wage to cover benefits, overhead, and short specialist engagements.","Only synthetic data and local or already available compute are used."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Partner-specific prototype with enforceable parser, state-space checker, result schema, audit record, access control, integration adapter, threat model, and small usability study.","confidence":"LOW","assumptions":["Two to five specialist person-months across software, formal methods, qualitative methods, security, and UX.","No migration of legacy workflows and no identifiable transcript corpus.","Existing platform authentication, storage, logging, and deployment infrastructure can be reused."],"source_ids":["S5","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production hardening, privacy and ethics review, platform integration, independent assurance, monitoring, incident and rollback procedures, documentation, training, and a controlled evaluation before any active-study use.","confidence":"LOW","assumptions":["Roughly two to five full-time-equivalent years of mixed engineering, assurance, methods, security, and product work.","The launch covers one platform and one narrow DSL, not arbitrary third-party scripting languages.","The organization already has lawful transcript storage and participant-rights workflows; building those systems from scratch is excluded."],"source_ids":["S5","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"One-platform maintenance: DSL/version review, proof and regression tests, security updates, governance review, user support, monitoring, incident response, and reclassification when language or event semantics change.","confidence":"LOW","assumptions":["Approximately 1.5-4 continuing FTE equivalents plus modest compute and external review.","Human substantive-method review remains necessary for UNKNOWN cases and is not treated as an oracle.","Volume is moderate and does not require a dedicated large-scale verification cluster."],"source_ids":["S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"External evidence strongly supports ambiguity and consequential stopping decisions in qualitative research, but the defining platform condition—arbitrary executable workflows plus a claimed universal Boolean validator—was not found.","source_ids":["S1","S2","S3","S5","S6"]},"externally_credible_adopter_or_authorizer":{"status":"UNCERTAIN","reason":"QDAcity is an identifiable adjacent platform, while researchers, reviewers, ethics boards, and funders are credible stakeholders. No source shows that any one of them wants, can integrate, or will authorize this specific intervention.","source_ids":["S2","S4","S5","S6"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim is narrow and testable through label-preservation, scope-enforcement, and counterexample tests against timeout Boolean, manual-review, and saturation-metric comparators.","source_ids":["S3","S5","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A two-week synthetic, non-deployment exercise with six predeclared workflows, explicit comparators, independent review, and clear falsifiers is bounded in time, data, authority, and cost.","source_ids":["S7","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The bounded step uses synthetic data, does not stop recruitment or certify substantive saturation, and can be rolled back by discarding the prototype. This gate does not clear deployment with participant data, which still requires privacy, ethics, and authority review.","source_ids":["S2","S4"]},"credible_cost_scope_and_range":{"status":"YES","reason":"Broad labor-equivalent bands, included and excluded scope, staffing assumptions, and uncertainty are explicit and anchored to 2025 national wage data. Vendor quotes and institution-specific overhead remain absent.","source_ids":["S8"]}},"next_evidence_step":"Run a 14-calendar-day, synthetic-data challenge with six frozen workflows: two safely terminating, one premature-stop path, one loop, one workflow simulating a supplied program, and one syntactically out of the finite fragment. Compare (A) a 30-second sandbox timeout converted to Boolean pass/fail, (B) manual SRQR-style methods review plus the Guest-style saturation calculation, and (C) the proposed finite checker/router. Pre-register expected outputs and require the router to emit EXACT only for fully enumerated in-fragment cases, BOUNDED when a depth limit is used, UNKNOWN for unresolved unrestricted cases, and OUT_OF_SCOPE for rejected syntax. An engineer drafts the source-to-workflow reduction; an independent formal-methods reviewer checks source-to-target direction, total computability of the translation, answer preservation, and the exact interaction semantics. Falsify advancement if any in-fragment expected result is wrong, the loop becomes false rather than UNKNOWN, the out-of-fragment case is certified, the reduction fails an obligation, the formal stopping predicate is judged by two independent qualitative methodologists not to represent the declared evidentiary condition, or the Boolean baseline preserves the same distinctions at materially lower complexity. Passing supports only a partnered prevalence and usability study, not deployment or an undecidability conclusion.","blocking_evidence":["No public evidence that a real qualitative platform accepts the declared class of arbitrary executable workflows and advertises a universal terminating Boolean validator.","No named platform owner, PI, ethics body, funder, or journal has expressed demand for the proposed computability-boundary intervention or committed to a pilot.","No independently reviewed, model-matched reduction or constructive decider exists for the candidate's exact workflow and transcript-interaction semantics.","No empirical evidence shows that researchers and reviewers correctly understand EXACT, BOUNDED, UNKNOWN, timeout, and OUT_OF_SCOPE labels.","No assessment covers privacy, participant withdrawal, retention, cross-border data handling, accessibility, or institutional legal requirements for a production integration.","Production cost bands lack architecture-specific estimates, vendor quotes, workflow volumes, and measured reviewer load.","Q-Sat AI is a preprint and its reported predictive performance was not independently validated in this review."],"research_disposition":"PROBLEM_PREVALENCE_STUDY","world_novelty_boundary":"No world-novelty conclusion is made. The bounded search found adjacent saturation metrics, computational decision support, qualitative reporting standards, and formal workflow soundness research, but did not search patents, non-indexed code, procurement records, private platform roadmaps, every language, or every jurisdiction. Patentability, freedom to operate, market size, realized impact, and global product novelty remain unmeasured.","arm":"PROPOSAL_FIRST","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":true,"material_progress_observed":true,"progress_targets":["Document at least one real platform's accepted workflow language, event model, current validator behavior, user-facing claims, and observed timeout or mislabel cases using platform-owned evidence.","Obtain a named platform owner or research-governance partner willing to authorize the synthetic challenge and, conditionally, a non-deployment pilot.","Complete independent review of the exact reduction or formally record why the computability status remains unresolved.","Measure guarantee-label comprehension and downstream decision behavior against a Boolean timeout baseline with qualitative researchers and reviewers.","Demonstrate semantic fidelity between the formal premature-stop property and a predeclared qualitative methodology; preserve substantive saturation as a human-authority decision.","Produce an architecture-specific privacy, security, participant-rights, ethics, and rollback assessment before identifiable data or active studies enter scope.","Replace labor-only deployment estimates with partner architecture estimates, expected workflow volumes, reviewer-load measurements, and at least one independent quote or bottom-up implementation plan."],"reason":"Web research resolves the general methodological need and exposes strong adjacent prior art, but it cannot establish the candidate's defining problem prevalence, adopter pull, semantic fidelity, label comprehension, or production authority. Those questions require platform access, interviews or observed workflows, independent proof work, and live comparative testing; further bounded browsing would not clear the uncertain gates."}}