Cohort Model¶
Recognize a spoken word incrementally by activating lexical candidates compatible with its onset, continuously selecting against mismatching competitors as acoustic evidence arrives, and integrating the surviving lexical interpretation with higher-level context.
Core Idea¶
The Cohort Model is an incremental account of spoken-word recognition. As the acoustic-phonetic input unfolds from a word's onset, it activates a cohort of lexical representations compatible with the evidence heard so far. New segments and timing information progressively distinguish candidates: incompatible words lose consideration or activation, better-fitting candidates survive, and one lexical interpretation becomes available for integration with syntactic and semantic context. Recognition is therefore not postponed until the word ends and not performed by an exhaustive serial dictionary search.
Marslen-Wilson and Welsh's 1978 experiments used shadowing and mispronunciation detection to study interactions between bottom-up speech analysis and contextual constraints. They proposed an active direct-access account in which early acoustic information activates lexical interpretations during continuous speech.[1] Marslen-Wilson's 1987 formulation names access, selection, and integration as the three basic functions and distinguishes two versions of the cohort model: an earlier partially interactive model and a revised account in which form-based access and selection are bottom-up while context contributes through integration.[2] Reference-grade treatment must preserve that evolution instead of presenting one simplified diagram as timeless doctrine.
The initial cohort is constrained by the word onset and the listener's lexical representations. If the input begins with a sequence compatible with captain, captive, caption, and capture, those items can be active competitors before later evidence separates them. A uniqueness point is the position at which the phonemic sequence corresponds to only one lexical item under a chosen lexicon and representation. It is a useful property of the input and lexicon, not necessarily the exact millisecond of behavioral recognition. Noise, coarticulation, frequency, lexical knowledge, segmentation, and decision criteria can move observed responses relative to that point.
Selection is graded in later interpretations. A candidate need not vanish after one idealized categorical mismatch in every experimental condition; goodness of fit, lexical frequency, acoustic uncertainty, and competition affect activation. The model nevertheless retains an onset-driven parallel candidate set and accumulating-evidence selection. Word frequency can influence activation or selection without becoming the model's defining identity. Context can facilitate integration and plausibility judgments while the 1987 bottom-up version denies that semantic context directly inserts an acoustically incompatible form into the initial cohort.
The model is primarily about spoken words. Extending it to visual word recognition requires a new input representation and comparison with models developed for orthographic processing. It is also not a general model of lexical retrieval in production, memory search, or sentence comprehension. The candidate survives catalog review because Selection alone does not specify onset activation, cohort membership, acoustic accumulation, or lexical integration; Word Frequency Effect is an empirical modulation; Phoneme and Phonology describe linguistic units, not an online recognition architecture.
Structural Signature¶
- The unfolding spoken input. Acoustic-phonetic evidence arrives over time rather than as a completed symbolic word.
- The mental lexical inventory. Stored word-form representations define which candidates can be activated.
- The onset-driven access rule. Early input activates multiple lexical items compatible with the initial evidence.
- The cohort. The currently active candidate population shares sufficient fit to the heard sequence.
- The accumulating evidence. Later segments, timing, and acoustic detail update candidate compatibility.
- The selection process. Mismatching or weaker competitors lose activation while the best-fitting representation gains relative support.
- The uniqueness or discrimination structure. Lexical overlap determines when the signal can in principle isolate a candidate.
- The integration process. Selected lexical syntactic and semantic information connects to the utterance context.
- The model-version boundary. Cohort I and revised Cohort II assign contextual influence differently.
- The behavioral evidence. Shadowing, gating, lexical decision, mispronunciation detection, eye movement, or neural measures test distinct predictions.
What It Is Not¶
- Not a model of group cohorts.
Cohortmeans a simultaneously active lexical candidate set. - Not visual word recognition by default. The foundational model concerns temporally unfolding speech.
- Not speech production. It models recognition/access from input rather than selecting a word to articulate.
- Not a serial exhaustive search. Multiple onset-compatible candidates are active in parallel.
- Not the uniqueness point alone. Lexical isolation is one property inside a broader access-selection-integration architecture.
- Not unrestricted top-down guessing. Revised versions keep form-based access and selection driven by speech evidence.
- Not the TRACE model. TRACE uses an interactive activation network with distinct feedback commitments.
Scope of Application¶
The Cohort Model is literal when spoken input activates an onset-compatible lexical candidate set whose membership or activation is updated online until selection and integration occur.
- Auditory lexical access. Explaining how word-form representations become active before acoustic offset.
- Lexical competition. Predicting effects of onset neighborhood and competitor overlap.
- Mispronunciation detection. Testing how mismatch position and sentence context affect online processing.
- Speech shadowing. Measuring rapid repetition and restoration during continuous speech.
- Gating paradigms. Presenting increasingly long word fragments to estimate candidate and recognition dynamics.
- Uniqueness-point analysis. Relating lexical inventory to the position where one candidate remains symbolically compatible.
- Context studies. Comparing partially interactive and bottom-up access/selection versions.
- Computational psycholinguistics. Implementing graded candidate activation while preserving testable architectural commitments.
Clarity¶
A clear use names the model version, language, participant population, lexical inventory, word segmentation assumptions, acoustic representation, candidate activation threshold, mismatch rule, frequency treatment, context locus, and behavioral task. Define cohort membership operationally and distinguish a phoneme-transcript uniqueness point from acoustic and behavioral recognition. Report onset competitors under the participant-relevant lexicon, not only a dictionary chosen for convenience. A gating response, eye fixation, shadowing latency, and neural signal index different processes. Context effects do not by themselves prove top-down alteration of lexical access; they may arise during selection, integration, or response. Model comparisons must state which architectural commitment, not merely which curve fit, differentiates them.
Manages Complexity¶
Continuous speech presents a combinatorial recognition problem because partial onsets match many words and segmentation is uncertain. The Cohort Model reduces it to a dynamically shrinking candidate population. Early parallel access preserves alternatives; accumulating acoustic evidence prevents commitment from depending on the whole lexicon indefinitely; selection and integration separate form matching from utterance-level interpretation. The simplification can become too categorical. Real speech contains coarticulation, reductions, noise, accents, and gradient similarity, and listeners differ in vocabulary. Modern implementations may therefore use activation strengths rather than a hard list. The abstraction remains useful if those refinements preserve onset-sensitive access and evidence-driven competition rather than turning cohort into any lexical-neighborhood model.
Abstract Reasoning¶
- Represent the incoming speech signal as temporally ordered acoustic-phonetic evidence.
- At onset, activate lexical forms whose beginnings are sufficiently compatible with the evidence.
- Maintain the active cohort rather than committing to the first plausible word.
- Update each candidate as new segmental, temporal, and acoustic detail arrives.
- Reduce or eliminate candidates whose form no longer fits under the declared mismatch rule.
- Allow lexical frequency and goodness of fit to influence competition without overriding the observed signal.
- Identify the best-supported lexical representation and distinguish this process from response execution.
- Integrate its syntactic and semantic information with sentence and discourse context.
- Compare predicted cohort dynamics with time-sensitive behavioral or neural measures.
- Revise the access, selection, or integration assumption implicated by a model-discriminating result.
Knowledge Transfer¶
The model transfers a general streaming-inference pattern: early evidence activates a population of compatible hypotheses; later evidence selects among them; higher-level interpretation uses the selected result. Transfer is strongest in domains with time-ordered signals and a finite candidate inventory. It is weaker where candidates do not share onset structure or where feedback changes the sensory evidence itself. The Cohort Model also transfers a methodological distinction between the informational point where one hypothesis becomes unique and the observed time at which an agent acts.
Examples¶
Canonical¶
A listener hears the beginning cap-. Several stored words may fit the onset. The cohort is activated before the sound sequence is complete. As the vowel, consonants, and timing continue, candidates incompatible with the signal lose activation. If the sequence becomes uniquely compatible with captain under the listener's lexicon, the form can be selected and its syntactic and semantic information integrated into the sentence. A plausible sentence context may speed integration, but under the revised bottom-up version it does not make a form survive an incompatible acoustic sequence.[2]
Mapped back: unfolding onset → parallel compatible word forms → acoustic mismatch-based competition → selected lexical entry → syntactic/semantic integration.
Applied / In Practice¶
In a mispronunciation-detection experiment, a word is altered early or late in a constraining sentence. The model predicts different opportunities for cohort activation and contextual integration across mismatch positions. Fluent restoration in shadowing is not direct proof that the acoustic mismatch was absent from access; the response may reflect integration or output repair. Marslen-Wilson and Welsh used the contrast between shadowing and explicit detection to separate dependencies on bottom-up input and contextual constraint.[1]
Mapped back: controlled acoustic mismatch + sentence context → access/selection/integration predictions → task-specific response measures → architectural inference.
Structural Tensions¶
- Early parallel access vs. efficient selection. Broad cohorts preserve possibilities but create competition. Diagnostic: How many onset-compatible candidates are measurably active?
- Categorical mismatch vs. gradient speech. Ideal phoneme strings simplify noisy acoustics. Diagnostic: Does candidate activation fall continuously with acoustic distance?
- Uniqueness point vs. recognition time. Lexical uniqueness is informational, behavior is process-dependent. Diagnostic: Is the reported point computed or experimentally observed?
- Bottom-up access vs. contextual facilitation. Context affects behavior without necessarily rewriting form access. Diagnostic: Which stage must context enter to explain the result?
- Stable lexicon vs. listener differences. Cohorts depend on stored words and pronunciations. Diagnostic: Is competitor structure participant-relevant?
- Word-onset alignment vs. continuous segmentation. Natural speech lacks guaranteed boundaries. Diagnostic: How is the onset located in continuous input?
- Historical model vs. generic competition label. Many models activate lexical neighbors. Diagnostic: Are access, selection, integration, and onset cohort commitments retained?
Structural–Framed Character¶
The structure is streaming speech input, onset-compatible lexical access, parallel cohort, accumulating mismatch evidence, selection, uniqueness structure, and contextual integration. The frame is language, listener lexicon, acoustic representation, segmentation, model version, task, and noise. Changing the frame alters cohort size and observed timing without erasing the architecture. Model history is part of the frame because Cohort I and Cohort II differ in where context may act.
Structural Core vs. Domain Accent¶
The transferable core is partial evidence → parallel hypothesis set → incremental evidence-based selection → higher-level integration. The domain accent is spoken words, acoustic phonetics, mental lexicon, onset cohorts, lexical frequency, uniqueness points, shadowing, and mispronunciation detection. Remove the accent and Selection remains; preserve it and the Cohort Model is a distinct psycholinguistic architecture.
Instantiates / Related Primes¶
Selection is the strict parent by specialization. A population of onset-compatible lexical candidates is available, accumulating acoustic evidence provides the selection basis, and mismatching competitors lose continuation while one interpretation is retained. Selection is broader and does not imply speech, lexical cohorts, or integration.
The prospective workspace queue contains one strict upward edge to prime:selection. No live DAG mutation is authorized.
Relationships to Other Abstractions¶
Current abstraction Cohort Model Domain-specific
Parents (1) — more general patterns this builds on
-
Cohort Model is a kind of Selection Prime
Selection is the strict parent by specialization.A population of onset-compatible lexical candidates is available, accumulating acoustic evidence provides the selection basis, and mismatching competitors lose continuation while one interpretation is retained. Selection is broader and does not imply speech, lexical cohorts, or integration. The prospective workspace queue contains one strict upward edge to
prime:selection. No live DAG mutation is authorized.
Hierarchy path (1) — routes to 1 parentless root
- Cohort Model → Selection
Neighborhood in Abstraction Space¶
Cohort Model sits in a sparse region of the domain-specific corpus (78th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Speech Planning & Lexical Perception (5 abstractions)
Nearest neighbors
- Underlying Representation — 0.85
- Lexical diffusion — 0.83
- Phonological Awareness — 0.83
- TRACE (psycholinguistics) — 0.82
- Word Frequency Effect — 0.82
Computed from structural-signature embeddings · 2026-09-08
Not to Be Confused With¶
- TRACE Model. Interactive-activation network with different feedback and representational commitments.
- Neighborhood Activation Model. Spoken-word framework emphasizing similarity neighborhoods under another architecture.
- Logogen Model. Threshold-based lexical recognition account that motivated contrasts in the original work.
- Serial Search Model. Examines lexical candidates sequentially rather than as an onset cohort.
- Uniqueness Point. Position where a lexical sequence becomes unique, not the entire model.
- Word Frequency Effect. Empirical processing advantage that can modulate candidates without defining the architecture.
- Lexical Retrieval in Production. Selecting a word to speak rather than recognizing incoming speech.
References¶
[1] William D. Marslen-Wilson and Alan Welsh, Processing Interactions and Lexical Access during Word Recognition in Continuous Speech, Cognitive Psychology 10, no. 1 (1978): 29–63, https://doi.org/10.1016/0010-0285(78)90018-X. registry ↩a ↩b
[2] William D. Marslen-Wilson, Functional Parallelism in Spoken Word-Recognition, Cognition 25, nos. 1–2 (1987): 71–102, https://doi.org/10.1016/0010-0277(87)90005-9. registry ↩a ↩b