Lexicostatistics¶
A comparative-linguistic method that quantifies shared basic-vocabulary cognates among languages to estimate degrees of lexical relationship.
Core Idea¶
Cognacy must be established rather than inferred from surface similarity, word-list meaning and borrowing controls matter, percentage similarity does not reconstruct a proto-language and glottochronological dating adds a controversial constant-rate assumption not required by lexicostatistics. Standardized concepts are elicited across languages, forms are coded into cognate classes and pairwise or multilateral shared-cognate proportions produce a distance or similarity matrix for comparison and clustering. The abstraction is therefore identified by a declared carrier, a transformation or constraint over that carrier, and an invariant that tells an analyst whether the named structure is genuinely present.
Scope of Application¶
Lexicostatistics belongs to comparative linguistics and is useful where the analyst can specify the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, then evaluate the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit.
Clarity¶
The abstraction clarifies a crowded vocabulary by making the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit the center of the account. A claim should name the carrier, the governing operation or relation, the applicable assumptions, and the recognition test.
Manages Complexity¶
Without the abstraction, an analyst must reason directly over many local details: the carrier roles, admissibility assumptions, competing conventions, derived invariants, boundary cases, and proof or validation obligations specific to Lexicostatistics. Lexicostatistics compresses them into the roles in the structural signature. That compression permits comparison across instances without erasing the variables that determine validity. It also exposes which details may be varied safely and which are constitutive.
Abstract Reasoning¶
- Identify the carrier. State what the elements, states, objects, or observations are: the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets. Reject examples whose alleged carrier belongs to a different problem. 2. Lock the constitutive rule. Express the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit independently of one notation or implementation.
Knowledge Transfer¶
Knowledge transfers strongly among subfields of comparative linguistics because they reuse the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, Standardized concepts are elicited across languages, forms are coded into cognate classes and pairwise or multilateral shared-cognate proportions produce a distance or similarity matrix for comparison and clustering., and type the carrier, state every parameter and convention in the definition, test that the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit, compare the nearest accepted identity, and report counterexamples, uncertainty, and limiting cases.
Relationships to Other Abstractions¶
Current abstraction Lexicostatistics Domain-specific
Parents (1) — more general patterns this builds on
-
Lexicostatistics is a kind of Measurement Prime
The proposed strict upward parent is
prime:measurement.
Hierarchy path (1) — routes to 1 parentless root
- Lexicostatistics → Measurement
Neighborhood in Abstraction Space¶
Lexicostatistics sits in a crowded region of the domain-specific corpus (26th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Historical & Quantitative Linguistics (14 abstractions)
Nearest neighbors
- Quantitative comparative linguistics — 0.93
- Lexicalization — 0.92
- Corpus linguistics — 0.91
- Codification (linguistics) — 0.90
- Componential analysis — 0.90
Computed from structural-signature embeddings · 2026-09-08