Skip to content

Lexicostatistics

A comparative-linguistic method that quantifies shared basic-vocabulary cognates among languages to estimate degrees of lexical relationship.

Version
v1 · 2026-09-08 · History
Domain-specific #
5312
Origin domain
comparative linguistics
Subdomain
comparative linguistics

Core Idea

Cognacy must be established rather than inferred from surface similarity, word-list meaning and borrowing controls matter, percentage similarity does not reconstruct a proto-language and glottochronological dating adds a controversial constant-rate assumption not required by lexicostatistics. Standardized concepts are elicited across languages, forms are coded into cognate classes and pairwise or multilateral shared-cognate proportions produce a distance or similarity matrix for comparison and clustering. The abstraction is therefore identified by a declared carrier, a transformation or constraint over that carrier, and an invariant that tells an analyst whether the named structure is genuinely present.

Scope of Application

Lexicostatistics belongs to comparative linguistics and is useful where the analyst can specify the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, then evaluate the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit.

Clarity

The abstraction clarifies a crowded vocabulary by making the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit the center of the account. A claim should name the carrier, the governing operation or relation, the applicable assumptions, and the recognition test.

Manages Complexity

Without the abstraction, an analyst must reason directly over many local details: the carrier roles, admissibility assumptions, competing conventions, derived invariants, boundary cases, and proof or validation obligations specific to Lexicostatistics. Lexicostatistics compresses them into the roles in the structural signature. That compression permits comparison across instances without erasing the variables that determine validity. It also exposes which details may be varied safely and which are constitutive.

Abstract Reasoning

  1. Identify the carrier. State what the elements, states, objects, or observations are: the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets. Reject examples whose alleged carrier belongs to a different problem. 2. Lock the constitutive rule. Express the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit independently of one notation or implementation.

Knowledge Transfer

Knowledge transfers strongly among subfields of comparative linguistics because they reuse the typed comparative linguistics carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, Standardized concepts are elicited across languages, forms are coded into cognate classes and pairwise or multilateral shared-cognate proportions produce a distance or similarity matrix for comparison and clustering., and type the carrier, state every parameter and convention in the definition, test that the languages and varieties, standardized basic-vocabulary concept list, lexical forms and transcription, cognate judgment method and coding, borrowing chance resemblance and missing-data treatment, shared-cognate numerator and comparable-item denominator, similarity or distance measure, clustering or network analysis, uncertainty and sensitivity and distinction from comparative reconstruction and glottochronology are explicit, compare the nearest accepted identity, and report counterexamples, uncertainty, and limiting cases.

Relationships to Other Abstractions

Local relationship map for LexicostatisticsParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.LexicostatisticsDOMAINPrime abstraction: Measurement — is a kind ofMeasurementPRIME

Current abstraction Lexicostatistics Domain-specific

Parents (1) — more general patterns this builds on

  • Lexicostatistics is a kind of Measurement Prime

    The proposed strict upward parent is prime:measurement.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Lexicostatistics sits in a crowded region of the domain-specific corpus (26th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Historical & Quantitative Linguistics (14 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08