Collostructional Analysis¶
A corpus method for measuring lexical preference for a grammatical-construction slot against a frequency baseline.
Core Idea¶
Collostructional analysis asks which words preferentially fill a slot in a specified grammatical construction. It compares corpus counts for a lexeme inside and outside that construction, or across two competing constructions, then ranks attraction, repulsion, or distinctiveness against a frequency baseline. The ranked words can inform an interpretation of constructional meaning; no single p-value supplies that meaning automatically.[ref-428d925da53e][ref-82b0aad5b481]
Scope of Application¶
For a single construction, the original research examined the noun slot of [N waiting to happen] and the verb slot of the into-causative. A distinctive-collexeme variant compared verbs in the ditransitive and to-dative patterns. These settings share construction-conditioned lexical counting but use different reference classes.[ref-428d925da53e][ref-82b0aad5b481]
Clarity¶
Frequent occurrence is not preferential occurrence. The method compares a word's in-slot count with what its overall distribution and the construction's frequency predict. A small Fisher p-value indicates a departure under the test's null; observed-versus-expected counts identify the direction of preference, and the lexical profile requires separate semantic interpretation.[ref-428d925da53e][ref-82b0aad5b481]
Manages Complexity¶
Repeated four-cell comparisons turn a long corpus concordance into a ranked set of candidate collexemes. That makes patterns inspectable without ignoring the choices that determine them: corpus coverage, lemma coding, construction boundaries, slot identification, and the chosen comparison class.[^ref-428d925da53e]
Abstract Reasoning¶
In the 2004 dative analysis, give occurred 461 times in the ditransitive and 146 in the to-dative; the authors' expected counts were about 213 and 394. The large positive deviation in the ditransitive, together with the test, supports a preference for that construction even though give appears in both. A raw 461-versus-146 comparison alone would miss the different construction totals.[^ref-82b0aad5b481]
Knowledge Transfer¶
The procedure can be reused for other grammatical construction types when their slots and tokens can be delimited. Generic word co-occurrence or nonlinguistic association is not collostructional analysis without the construction-and-slot carrier.
[^ref-428d925da53e]: Anatol Stefanowitsch and Stefan Th. Gries, “Collostructions: Investigating the interaction of words and constructions”, International Journal of Corpus Linguistics 8, no. 2 (2003), 209–243, especially §2.2 and §3.2.1. [^ref-82b0aad5b481]: Stefan Th. Gries and Anatol Stefanowitsch, “Extending collostructional analysis: A corpus-based perspective on ‘alternations’”, International Journal of Corpus Linguistics 9, no. 1 (2004), 97–129, especially §2 and Table 1.
Neighborhood in Abstraction Space¶
Collostructional Analysis sits in a moderately populated region (57th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Language, Mind & Meaning-Making (57 abstractions)
Nearest neighbors
- Tagmeme — 0.86
- Factored Language Model — 0.85
- Kneser–Ney Smoothing — 0.85
- Poetic Metre — 0.85
- Topic Marker — 0.85
Computed from structural-signature embeddings · 2026-10-08