Skip to content

Extended Boolean model

An information-retrieval model that relaxes exact Boolean matching by incorporating partial matching and term weights.

Version
v2 · 2026-09-06 · History
Domain-specific #
1813
Origin domain
information retrieval
Subdomain
graded Boolean retrieval
Aliases
Extended Boolean retrieval model, P-norm model

Core Idea

Extended Boolean model is an information-retrieval model that relaxes exact Boolean matching by incorporating partial matching and term weights. [1]

The extended Boolean model preserves an AND/OR query tree while replacing exact binary satisfaction with graded term weights and p-norm aggregation. Documents receive similarity scores and can be ranked; the p parameter controls how strictly each connective approximates classical Boolean logic.

Its operative boundary is not supplied by the name alone. Preserve this identity: An information-retrieval model that relaxes exact Boolean matching by incorporating partial matching and term weights. Validity boundary: Retrieval scores must combine Boolean query structure with graded term weighting under the extended model; pure vector-space ranking or exact Boolean matching is different. The entry therefore captures a reusable specialist role structure rather than a topic label, a single historical instance, or a loose analogy.

Structural Signature

Sig role-phrases:

  • the document representation — weighted term coordinates for each indexed document
  • the weighted query terms — term importance values in the request
  • the Boolean structure — the AND/OR expression tree retained from the query
  • the p-norm operator — the graded aggregation rule for each connective
  • the strictness parameter — p controlling movement from soft compensation toward Boolean behavior
  • the similarity score — a graded document–query match value
  • the ranked output — documents ordered by score rather than only accepted or rejected
  • the normalization — scaling that keeps term weights and scores comparable

Recognition test. A case qualifies only when the analyst can map the declared the document representation, the weighted query terms, the Boolean structure, the p-norm operator, the strictness parameter and preserve the specialist validity conditions. Shared vocabulary, a similar output, or a generic instance of one parent relation is insufficient.

What It Is Not

  • Not exact Boolean retrieval. Exact retrieval produces a binary set without partial-match ranking.
  • Not the vector-space model. The Boolean expression tree remains semantically active.
  • Not fuzzy logic in general. This is a specific p-norm information-retrieval construction.
  • Not probabilistic retrieval. Scores are not necessarily relevance probabilities.
  • Not arbitrary weighted keyword search. The model specifies how weights compose under AND and OR.

Scope of Application

The abstraction recurs literally within text retrieval systems that need Boolean query structure together with partial matching and ranked results. The following habitats preserve the same recognition machinery; they are not invitations to extend the name metaphorically.

  • Soft conjunction. documents missing one term can receive penalized nonzero scores.
  • Soft disjunction. multiple matched terms increase score without exact set union only.
  • Weighted queries. term importance influences branch aggregation.
  • Ranked Boolean interfaces. users retain explicit connectives but receive an ordering.
  • Parameter studies. p tunes strictness against recall and ranking quality.

Clarity

Specify term-weight normalization, the exact AND and OR formulas, p values, treatment of nested operators, and score direction. 'Boolean plus ranking' is too broad; different fuzzy and probabilistic models yield different compensation and semantics.

A practical identification audit begins with the typed roles rather than the title: establish the document representation, verify the weighted query terms, then test the remaining conditions and exclusions. If the case retains only the portable skeleton described below, it should be named through a parent abstraction rather than as Extended Boolean model.

Manages Complexity

The model interpolates between interpretable query logic and graded retrieval. One expression can express structural requirements while avoiding the brittle all-or-nothing boundary of classical Boolean matching.

The compression remains accountable because each simplification has a named failure condition. Disagreement can be localized to a missing role, an invalid assumption, an ambiguous measurement, or a neighboring abstraction instead of being hidden inside an unanalyzed label.

Abstract Reasoning

R1. Represent documents and query terms on a compatible normalized weight scale. R2. Parse the query into an explicit Boolean tree. R3. Evaluate leaves and combine them with the declared p-norm AND/OR formulas. R4. Propagate scores through nesting and rank documents. R5. Validate p and weighting choices against relevance judgments and Boolean boundary cases.

These moves separate definition, derivation, measurement, and interpretation. A formal consequence does not by itself prove that an observed case instantiates the abstraction, while an observed resemblance does not relax the formal or institutional recognition conditions.

Knowledge Transfer

The model transfers literally to weighted document retrieval with p-norm Boolean composition. Compositionality and similarity measure are parents; general fuzzy rules or hybrid filters are not this model without its formulas.

The transfer boundary is explicit: DOMAIN-SPECIFIC PASS / PRIME FAIL: The model applies across weighted queries and document collections where strict Boolean result sets are too coarse. Literal recognition retains the specialist vocabulary and validity conditions of information retrieval; outside that setting only broader parent operations transfer. The safe move beyond the home habitat is to carry the applicable parent relation and leave the specialist name behind unless every defining role remains literal.

Examples

Canonical: a soft AND query

For 'renewable AND storage,' a document strong on both terms ranks highest, while one weak on storage is penalized rather than categorically excluded. Increasing p makes the conjunction behave more like the minimum-like strict Boolean requirement. [1]

Mapped back: the document representation; the weighted query terms; the Boolean structure; the p-norm operator; the strictness parameter; the similarity score.

Applied / In Practice: nested weighted retrieval

A query combines a high-weight phrase branch with an OR branch of synonyms. Each subtree is scored under its declared extended operator, then the root score ranks documents while retaining which query structure produced the result. [2]

Mapped back: the weighted query terms; the Boolean structure; the p-norm operator; the ranked output; the normalization.

Structural Tensions

T1: Boolean interpretability vs partial compensation. Soft matching improves recall but can return a document violating an expected hard requirement. Diagnostic: Which operators are truly mandatory?

T2: p flexibility vs tuning burden. Different strictness values reshape rankings. Diagnostic: How is p validated and exposed?

T3: Term weights vs connective semantics. A large weight can dominate a branch in ways users do not expect. Diagnostic: Are score explanations available?

T4: Nested structure vs normalization. Subtree scale differences can distort root aggregation. Diagnostic: Are branches comparable?

T5: Similarity score vs relevance probability. A normalized geometric score may look probabilistic without calibration. Diagnostic: How is the output interpreted?

T6: Domain autonomy vs prime reduction. Compositionality and Similarity Measure omit the specialist objects, constraints, and validity tests named above. Diagnostic: Would retaining only the portable parent pattern still satisfy the recognition test?

Structural–Framed Character

The five-criterion aggregate is 0.15 (structural). The judgment is criterion-specific:

  • Vocabulary travels — low (0.25). The complete vocabulary remains tied to the typed roles in the Structural Signature.
  • Evaluative weight — low (0.00). Application carries the stated degree of normative or interpretive judgment beyond structural recognition.
  • Institutional origin — low (0.25). The abstraction depends to this degree on a scholarly, technical, legal, or social convention.
  • Human-practice bound — low (0.00). Recognition depends to this degree on organized practice, language, measurement, or institutional action.
  • Import versus recognize — low (0.25). Beyond its home habitat, use of the full name increasingly becomes analogy rather than literal recognition.

The portable skeleton is logical composition is relaxed into graded similarity while retaining the query's explicit operator structure. The named abstraction remains structural because that skeleton alone does not supply its specialist objects, constraints, or tests.

Structural Core vs. Domain Accent

Structural core: Logical composition is relaxed into graded similarity while retaining the query's explicit operator structure.

Domain accent: Document term weights, boolean query trees, p-norm and and or, strictness parameters, partial matching, normalized scores, and ranked retrieval.

Why it does not clear the prime bar: Compositionality and similarity travel; the extended Boolean model is their specific p-norm retrieval formulation. Generalization therefore routes through parent abstractions; preserving the specialist name requires the full accent.

  • Compositionality (prime:compositionality). A query score is built recursively from term and connective meanings.
  • Similarity Measure (prime:similarity_measure). Documents receive graded match scores used for ranking.

These are prose placement proposals only. They create no dag_edges; endpoint, redundancy, and cycle checks are recorded separately in the bundle's placement memo.

Relationships to Other Abstractions

Local relationship map for Extended Boolean modelParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.ExtendedBoolean modelDOMAINPrime abstraction: Similarity Measure — presupposesSimilarityMeasurePRIMEPrime abstraction: Compositionality — is a kind ofCompositionalityPRIME

Current abstraction Extended Boolean model Domain-specific

Parents (2) — more general patterns this builds on

  • Extended Boolean model is a kind of Compositionality Prime

    Compositionality (prime:compositionality).

  • Extended Boolean model presupposes Similarity Measure Prime

    Similarity Measure (prime:similarity_measure).

Hierarchy paths (3) — routes to 3 parentless roots

Neighborhood in Abstraction Space

Extended Boolean model sits in a sparse region of the domain-specific corpus (76th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (1565 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08

Not to Be Confused With

  • Boolean retrieval. exact set operations with binary matching. Tell: Can partial matches receive scores?
  • Vector-space model. cosine or related similarity without required Boolean structure. Tell: Does an AND/OR tree govern composition?
  • Fuzzy retrieval. a broader family using fuzzy sets or logics. Tell: Are Salton–Fox–Wu p-norm operators used?
  • Probabilistic retrieval. ranking by estimated relevance probabilities. Tell: Is the score geometric or probabilistic?
  • BM25. a term-saturation ranking function. Tell: Are query connectives explicitly composed?

References

[1] Gerard Salton, Edward A. Fox, and Harry Wu, “Extended Boolean Information Retrieval”, Communications of the ACM 26(11) (1983), 1022–1036. registry ↩a ↩b

[2] Gerard Salton, “The Use of Extended Boolean Logic in Information Retrieval”, SIGMOD Record 14(2) (1984), 277–285. registry