Skip to content

Literature-Based Discovery

Generate testable hypotheses by linking complementary relations stated in separate scholarly literatures—often A–B and B–C—to propose an unstated A–C connection.

Version
v1 · 2026-08-30 · History
Domain-specific #
2196
Origin domain
information science
Subdomain
knowledge discovery from literature
Aliases
Literature-related discovery, LBD, Swanson linking

Core Idea

Literature-based discovery (LBD) is the information-science practice of generating hypotheses by connecting claims that are explicit in separate scholarly literatures but whose joint implication has not been explicitly investigated. Its canonical ABC pattern begins with an established relation between concept A and intermediary B and another between B and C. If the A and C literatures are substantially disconnected, the shared B suggests a potentially novel A–C relation for expert assessment and empirical testing.

Don Swanson pioneered the method in the 1980s. His best-known case joined literature about Raynaud disease with literature about fish oil through intermediate concepts such as blood viscosity and platelet aggregation. The proposed therapeutic relation was later investigated prospectively.

Scope of Application

LBD is used in biomedical informatics, drug repurposing, adverse-event detection, gene–disease association, biomarker discovery, disease mechanism studies, research-policy analysis, materials and environmental research, and interdisciplinary collaboration discovery. It is strongest where a large indexed literature has good entity normalization and partially complementary research communities.

The method can incorporate curated databases alongside text if provenance distinguishes sources. Semantic typing can restrict paths—for example disease–process–drug rather than arbitrary word chains. Contextualized relations improve on raw co-occurrence by preserving direction, negation, species, experimental setting, and causal role.

Clarity

In open discovery, choose A, retrieve its associated Bs, then expand each B to candidate Cs outside the starting literature. Rank Cs and display the strongest paths. In closed discovery, choose A and C and search for Bs that connect them, producing possible mechanisms or evidentiary bridges.

Manages Complexity

Scientific specialization distributes relevant facts across journals, vocabularies, and communities. No researcher can read every adjacent literature. LBD turns this fragmentation into a search space: normalized concepts become nodes, explicit relations become edges, and unexamined paths become candidates.

The ABC abstraction is a powerful compression. It reduces millions of documents to interpretable bridge patterns, while filters and rankings control path explosion. Evidence chains retain human auditability that an opaque similarity score lacks.

Abstract Reasoning

  1. Literature separation can hide public knowledge. Facts may be individually published yet jointly unrecognized. 2. Bridge diversity raises robustness. Several mechanistically distinct Bs reduce dependence on one extraction error. 3. Typed relations outperform blind transitivity. Valid inference depends on predicate semantics and direction. 4. Novelty is time-indexed. A candidate can be a genuine pre-cutoff prediction even if published later. 5. Ranking creates selection bias. Benchmarking only famous rediscoveries can reward systems tuned to a tiny canon.

Knowledge Transfer

The method transfers across fields when documents can be normalized into entities and relations. Its workflow—retrieve, normalize, link, rank, inspect, validate—remains stable even when domain ontologies change.

Transfer is weaker in fields where claims depend on long arguments, images, tacit practices, or concepts not captured as simple relations. Critics correctly note that science is not exhausted by ABC triples. Richer graph, embedding, and analogy systems extend the method but must preserve evidentiary traceability.

Relationships to Other Abstractions

Local relationship map for Literature-Based DiscoveryParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Literature-BasedDiscoveryDOMAINPrime abstraction: Abductive Reasoning — is part ofAbductiveReasoningPRIME

Current abstraction Literature-Based Discovery Domain-specific

Parents (1) — more general patterns this builds on

  • Literature-Based Discovery is part of Abductive Reasoning Prime

    multiple independent paths strengthen a candidate.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Literature-Based Discovery sits in a sparse region of the domain-specific corpus (96th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (1565 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08