Skip to content

Constraint Grammar

A rule-based language-analysis method that uses contextual constraints to disambiguate token readings and assign linguistic tags.

Version
v1 · 2026-09-28 · History
Domain-specific #
8668
Domain group
Humanities
Origin domain
Linguistics & Semiotics
Subdomains
Computational Linguistics, Morphosyntactic Disambiguation → Linguistics & Semiotics
Aliases
CG

Core Idea

Constraint Grammar describes a way to analyze running language as a sequence of token cohorts: each token arrives with one or more candidate readings, and contextual rules revise those readings or attach tags. Operations include selection, removal, addition, and replacement. Context may be local or distant and may involve linked or negated tests. The load-bearing identity is the authored linguistic rule operating over candidate analyses, not the word 'constraint' by itself.

A CG system can progress from safer disambiguation rules to more heuristic and syntactic annotation stages. Rules must not remove the last reading of a kind, a safeguard against an empty analysis. A CG-3 rule for 'was' demonstrates the mechanism directly. This is distinct from the morphological analyzer that proposes readings, a purely statistical tagger, or a generic solver. Reported accuracy figures and language coverage describe implementations; they are not guaranteed by membership in the method.

How would you explain it like I'm…

Crossing Out Wrong Word Meanings

Some words can mean more than one thing, like 'bark' can be a dog's sound or the outside of a tree. Constraint Grammar is a set of rules people write for a computer that look at the words nearby and cross out the meanings that don't fit. There's one safety rule: never cross out a word's very last meaning, so every word keeps at least one.

Rules That Pick the Right Meaning

Constraint Grammar is a way to help computers understand sentences. Each word starts with a list of possible labels, like 'noun' or 'verb.' People write rules that look at nearby or even far-away words and then pick a label, remove a label, or add a new tag. The rules usually start careful and safe, then get bolder. One safety rule: a word's last remaining label of a kind can't be removed, so no word is left with nothing.

Rule-Based Reading Disambiguation

Constraint Grammar (CG) analyzes running text as a series of tokens, each arriving with a 'cohort' of possible readings, such as different parts of speech or word forms. Hand-written contextual rules then act on these readings: they can select one, remove some, add tags, or replace them. The rules can test nearby or distant words, link several tests together, or require that something is absent. Systems usually run safer disambiguation rules first and more heuristic ones and syntactic labeling later, and rules may never delete a word's last reading. CG is different from the tool that proposes the readings in the first place and from a purely statistical tagger that learns patterns without authored rules.

 

Constraint Grammar is a rule-based framework for analyzing running language as a sequence of token cohorts, each token carrying one or more candidate readings supplied by a morphological analyzer. Contextual rules written by linguists revise the cohorts through operations that select, remove, add or replace readings and tags. Contexts can be local or long-distance and can include linked tests and negated conditions. Grammars are typically ordered from safer disambiguation rules toward more heuristic rules and syntactic annotation. A key safeguard forbids removing the last reading of a kind, so no token ends up with an empty analysis. The defining feature is authored linguistic rules operating over candidate analyses; it is distinct from the analyzer that proposes readings, from purely statistical taggers, and from generic constraint solvers. Accuracy figures and language coverage belong to particular implementations, not to the method itself.

Structural Signature

Sig role-phrases:

  • analyzed token cohort — Carries a surface token and its candidate linguistic readings. It is constitutive. Counterfactual: Unanalyzed raw text has not yet supplied the reading set on which CG rules operate.
  • linguistic readings and tags — Represent alternative morphology, syntax, and other annotations to be revised. It is constitutive. Counterfactual: A rule that changes arbitrary characters without linguistic readings is not this method.
  • context-conditioned rule — Tests surrounding tags or words before a select, remove, add, or replace action. It is constitutive. Counterfactual: A context-free lexicon lookup alone lacks the contextual disambiguation relation.
  • ordered rule application — Applies rule sets in stages so cautious decisions can precede heuristic ones. It is operating condition. Counterfactual: The source describes staged application, not a claim that one isolated rule is a complete CG parser.
  • retained analysis — Leaves one or more supported readings and possibly later syntactic labels for downstream processing. It is boundary. Counterfactual: Erasing every reading would destroy the token analysis rather than disambiguate it.

What It Is Not

  • Any formal grammar. Production rules that generate sentences do not automatically disambiguate token cohorts.
  • Morphological analysis alone. Producing possible readings supplies CG's input but does not apply contextual selection.
  • A statistical tagger by label. Probabilistic scores without authored context-action rules do not instantiate the described paradigm.
  • Guaranteed accuracy. Published system scores do not follow from the abstraction for every language or corpus.
  • Closest near-miss. A finite-state morphological analyzer may generate the initial alternatives used by CG, but generation alone does not perform CG's contextual selection and mapping.

Scope of Application

  • Morphosyntactic annotation. Track how ambiguous word analyses are pruned or labeled by contextual evidence.
  • Parser comparison. Separate CG rule action from a generative grammar or a learned tagger.
  • Language-tool integration. Locate a CG stage between morphological analysis and downstream translation or treebank use.
  • Error analysis. Identify whether a wrong reading came from an input analyzer, a rule condition, or rule ordering.

Clarity

First locate the cohort and its alternative readings. Then identify the condition being tested in surrounding text and the operation applied to those readings. A lexicon may supply alternatives; CG decides among or augments them under context. Do not confuse the presence of a rule language with proof that one implementation reaches its reported evaluation score.

Manages Complexity

CG compresses many local linguistic choices into reusable condition-action rules, yet interactions among stages and far-reaching conditions still matter. Its staged approach can preserve unresolved ambiguity instead of guessing whenever evidence is weak. A description that omits the cohort, rule action, or retained reading hides exactly where an analysis changed.

Abstract Reasoning

  1. List the candidate readings for the token rather than naming only its surface form.
  2. State the contextual evidence that the rule tests.
  3. Name whether the rule selects, removes, adds, or replaces an annotation.
  4. Check the resulting reading set and whether a later stage supplies syntactic labels.
  5. Distinguish input analyzer errors and downstream-use claims from the CG operation itself.

Knowledge Transfer

The cohort–context–rule–retained-reading pattern transfers across languages and CG implementations when each language supplies suitable analyses and authored linguistic conditions. A specific English pronoun rule, tag inventory, or published score does not transfer unchanged to another language or corpus.

Examples

Canonical

A VISL CG-3 rule removes the first-person past-tense reading of 'was' when its context rejects a first-person pronoun to the left; the surviving token reading is third person. This illustrates disambiguation on an actual cohort, not a generic grammar slogan.

Mapped back: analyzed token cohort → the 'was' cohort; linguistic readings and tags → past-tense first- and third-person alternatives; context-conditioned rule → REMOVE (verb p1) with left-context test; ordered rule application → a rule in a staged CG-3 grammar; retained analysis → third-person reading remains.

Applied / In Practice

Apertium uses converted Constraint Grammars for Celtic and other language-processing systems. There, authored contextual rules act on morphological cohorts as a real translation-pipeline analysis component; this does not establish a particular accuracy for those systems.

Mapped back: analyzed token cohort → analyzed source-language tokens; linguistic readings and tags → morphological alternatives; context-conditioned rule → Apertium-deployed converted CG rules; ordered rule application → grammar stage before later language processing; retained analysis → selected readings supplied onward.

Structural Tensions

T1 — Safe Disambiguation versus Coverage. Cautious rules preserve plausible readings, while broader heuristics can resolve more ambiguity at the risk of error.

Diagnostic: Which stage permits a heuristic, and what ambiguity remains?

T2 — Local Context versus Global Dependency. Rules can inspect nearby or unbounded contexts, increasing expressive reach but also interaction complexity.

Diagnostic: Does the linguistic decision genuinely require a nonlocal condition?

Structural–Framed Character

A provisional portable skeleton is contextual narrowing of alternative analyses by explicit rules. Constraint Grammar uses linguist-authored predicates to select, remove, add, or replace readings on token cohorts without indiscriminately deleting all options. One constraint is an ingredient, not the method's genus.

Evaluative weight: Analysis quality is empirical and language-dependent, not guaranteed by rules. Human-practice-bound: High, because linguists write tags, contexts, and rule order. Institutional origin: NLP practice developed implementations; no one English rule is universal. Vocabulary travels: The cohort–context pattern spans languages after rebuilding tag inventories. Import versus recognize: Recognize CG by token readings and staged contextual actions; generic constraint solving imports a different carrier.

Its character: An authored NLP method with portable ambiguity-reduction logic and linguistic token semantics.

Structural Core vs. Domain Accent

Skeletal core. Explicit contextual rules progressively narrow competing interpretations.

Domain-bound accent. Analyzed text tokens, candidate tags or readings, linguist-authored contexts, and staged rule actions define Constraint Grammar.

Why not prime. Constraint use is broad; context-free lookup or unrelated solver constraints lack this method.

  • Related — constraint. CG uses exclusion and admissibility conditions on readings, but a whole grammar pipeline is a method comprising many such constraints, not one constraint object.

  • Related — classification. Selecting a linguistic tag can classify a token; CG additionally specifies cohort alternatives, context conditions, and staged rule actions.

Neighborhood in Abstraction Space

Constraint Grammar sits in a crowded region of the domain-specific corpus (27th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Language Structure & Grammar Formalisms (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • Context-free grammar. Tell: Are rules generating structures, or revising readings in running text under context?
  • Morphological analyzer. Tell: Does it propose candidate readings or select among them?
  • Statistical POS tagger. Tell: Are decisions carried by authored context-action rules or learned probabilities?
  • General constraint solver. Tell: Are the constrained objects linguistic token readings?

References

  • Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Constraint_grammar (revision 1366126233).
  • Preserved source candidate: https://sourceforge.net/projects/vislcg/
  • Preserved source candidate: https://visl.sdu.dk/cg3ide.html
  • Preserved source candidate: https://visl.sdu.dk/svn/visl/tools/vislcg3/trunk/emacs/cg.el
  • Preserved source candidate: http://beta.visl.sdu.dk/cg3.html
  • Preserved source candidate: https://web.archive.org/web/20110822024405/http://giellatekno.uit.no/english.html
  • Preserved source candidate: https://web.archive.org/web/20060719144813/http://www.divvun.no/doc/tools/docu-sme-manual.html
  • Preserved source candidate: https://web.archive.org/web/20110722011002/https://victorio.uit.no/langtech/trunk/kt/fin/src/fin-dis.cg1
  • Preserved source candidate: https://archive.today/20110722011041/https://victorio.uit.no/langtech/trunk/kt/fin/src/fin-dis.rle

The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.