Co-citation¶
A relation between two cited works measured by how many other works cite both within a specified corpus.
Core Idea¶
Co-citation relates two cited documents when a third document cites both. Its raw frequency is the number of distinct citing documents that do so in a specified corpus. If C(a) and C(b) are the sets of corpus documents citing focal works a and b, respectively, the frequency is |C(a) ∩ C(b)|. Henry Small's original formulation explicitly used this common-citer intersection and distinguished the raw count from a possible relative frequency. A positive count establishes an observable joint-citation relation; a zero count means no joint citer is present in that declared corpus.[1]
The direction of the evidence matters. Co-citation looks inward toward the two focal works from later works that cite them. It is therefore capable of changing as a corpus or observation window admits new citing works. This contrasts with bibliographic coupling, which compares the focal works' own outgoing reference lists: adding a new later citer cannot change which sources those two already cite. Neither measure by itself states that the focal works agree, endorse one another, or have identical content. The trace is that another author used them together.[1][2]
Structural Signature¶
Sig role-phrases: cited pair → declared citing corpus → joint-citer test → count of distinct joint citers → bounded interpretation.
- Focal cited pair. Choose two distinct works
aandb. They are the relata. The pair, not any single work's citation total, is the unit being characterized.[1] - Declared citing corpus. Fix which later or otherwise eligible documents and which snapshot of their reference lists are searched. An index with different coverage can yield a different count for the same pair; the corpus is part of a reproducible claim.[1][3]
- Joint-citer predicate. A corpus document contributes only if its references include both
aandb. One-sided citations do not satisfy the predicate. The relevant lists areC(a)andC(b), not the references authored byaandb.[1] - Raw frequency. Count the distinct members of
C(a) ∩ C(b). The result is a nonnegative integer. A positive/zero binary edge or a weighted network can be derived from it, but normalization and clustering are subsequent choices rather than parts of the raw definition.[1] - Interpretive boundary. A joint reference is evidence of use in a common citing document. It does not encode the citing author's purpose or the contents of either work. Method papers, for example, may bridge mapped specialties without making their authors the intellectual leaders of each specialty.[3]
What It Is Not¶
- Not bibliographic coupling. In coupling, the two focal works are connected because they themselves cite a common earlier reference. In co-citation, a third work cites the two focal works. One pair may satisfy either, both or neither test. Reversing the citation arrows changes the relation.[1][2]
- Not direct citation. A citation from
atobis a one-way edge between them. It does not alone supply a third document citing both. Small reported a general resemblance of some co-citation patterns to direct-citation patterns, not an identity of the two measures.[1] - Not a proof of agreement or common subject. A citing work can juxtapose competing theories or use two methods for separate purposes. The common-citer event remains true while the interpretive inference fails.
- Not a required normalized score, cluster or map. Small proposed a relative frequency using an intersection-over-union denominator. That quotient answers a different comparison question; the raw frequency remains the number of joint citers. A graph layout or cluster analysis consumes the pairwise values but does not define them.[1]
- Not automatically author, journal or web-page co-citation. Those can be adapted relations if the focal unit and citing act are specified afresh. This entry's literal unit is the document in a scholarly citing corpus.
Scope of Application¶
In bibliometrics, the focal works are articles or other citable documents and the corpus is an indexed set of works with reference lists. The relation can be computed for one pair or for many pairs. Small's original paper gave a particle-physics example and proposed networks of co-cited papers for examining a specialty. The same document-level test can be repeated for another scientific specialty, provided the focal unit and corpus are consistently defined.[1]
In science mapping, a matrix of pairwise frequencies can be used to draw a network or find clusters. This is a downstream analytical setting: a cluster is a hypothesis about literature structure, not an extra condition each co-citation must meet. The underlying count is shaped by indexing coverage, time window and heavily cited methods papers; Small and Griffith's later work makes such coverage and interpretation limits salient.[1][3]
The entry does not specify a universal normalization, citation-context weighting or clustering algorithm. A fixed frozen corpus gives a fixed raw count; adding a joint citer to that corpus can increase it. Counting web hyperlinks or named authors would require a new explicit choice of nodes and evidence rule rather than silent transfer of the document definition.
Clarity¶
Co-citation makes “these papers are related” checkable by asking who cites them together? Suppose a and b share three topical keywords but no work in the chosen corpus cites both. Their topical similarity does not turn their co-citation count positive. Conversely, a joint citer may reference two opposed papers. The relation is clear because its membership test is a citation event, not an inferred semantic equivalence.[1]
It also resolves a frequently confusing duality. For bibliographic coupling, inspect the bibliographies inside a and b; for co-citation, inspect the works outside them that list both. The former's shared-reference count is fixed for the published bibliographies, whereas the latter can rise as new work cites the pair. Changing an index's coverage may change the observed latter count without changing the original documents.[1][2]
Manages Complexity¶
A field may contain too many citing documents to compare one by one. Co-citation compresses their reference-list overlaps into one count for each chosen pair. Across a set of cited papers, those counts form a matrix; a thresholded or weighted graph can then reveal concentrations of joint use. The compression is useful because the same explicit pair rule applies repeatedly, even as the number of source documents grows.[1]
The compression discards context: the count does not preserve which paragraph made the citation, whether the pair was contrasted, or why a methods paper is widely used. A normalized score may improve comparisons across unequally cited works, but it changes the numerical object and introduces denominator choices. Small's possible relative frequency is therefore a declared transformation, not a substitute silently slipped into the raw count.[1][3]
Abstract Reasoning¶
For a particular pair, form the set of citing works for each member and intersect them. If the intersection contains n distinct documents, the raw co-citation frequency is n. An included work citing only a cannot increase the pair's frequency; a newly included work citing both a and b increases it by one, assuming it is not already counted. The inference follows from set membership, not an estimate of thematic similarity.[1]
For many focal pairs, repeat the same calculation in one consistently declared corpus. A dense set of positive pair links may motivate a specialty map, but the analyst must separately ask whether the map reflects shared conceptual work, a common tool or citation practices. That second question requires evidence beyond the co-citation counts. The distinction is the point of the abstraction: a defensible relational trace can support exploration without bearing a stronger interpretation than its rule licenses.[1][3]
Knowledge Transfer¶
Within information science, the same document-level rule transfers from one indexed specialty to another: choose a cited pair, choose a citing corpus, test each citing work for both references, and count. Small's particle-physics example illustrates one such setting; a second literature can use the identical rule without needing the same cluster structure or citation density.[1]
At a higher level, the live Relation carries the portable idea of relata and an inclusion predicate, and Network can organize many accepted pair links. That does not make co-citation itself a prime or make any two jointly mentioned entities “co-cited.” Literal transfer to authors, journals or web pages would require declaring a different citation or linking event and reevaluating what the resulting count means. The named abstraction stays anchored in bibliometric document evidence.
Examples¶
Worked positive case. In a declared citing corpus {X,Y,Z,W}, X and Y both cite a and b, Z cites only a, and W cites neither. Then C(a) ∩ C(b)={X,Y} and the raw co-citation frequency of the pair is 2. If a fifth work citing both joins the corpus, it becomes 3. Mapped back: a,b are the focal cited pair; {X,Y,Z,W} is the corpus at the first snapshot; the joint-citer predicate selects X,Y; the count is 2; neither their motives nor the pair's semantic agreement follow from this arithmetic. This is an illustrative deduction from Small's rule, not a reported historical dataset.[1]
Original research application. Small described constructing a network of co-cited scientific papers and drew an example from particle physics. Mapped back: two cited physics papers are a focal pair; the indexed citing papers are the declared corpus; each common citer supplies one joint event; the pairwise frequency weights a potential network link; interpreting a network as specialty structure is a downstream analytical step. The original publisher abstract supports the setting and operation but does not warrant a numerical claim about a particular pair, so none is supplied here.[1]
Boundary case. Let a and b both cite earlier reference r, but suppose no work in the declared later corpus cites both. They are bibliographically coupled through r, yet their raw co-citation frequency is 0. The missing co-citation role is a joint citing document; sharing an outgoing reference is not a replacement.[1][2]
Structural Tensions¶
Raw observability versus normalized comparison. The unadjusted integer tells exactly how many joint citers were observed, but heavily cited works or uneven corpus coverage can make raw values hard to compare. A relative score can temper that exposure; it also changes the scale and may change rankings. Neither objective simply dominates the other. Diagnostic: Is the present question “how many works cite this pair together?” or “how strong is this pair relative to citation opportunity?” Small's raw and possible relative frequencies answer those differently.[1]
Readable map versus faithful citation context. Pairwise counts and clusters condense a sprawling literature into a tractable picture, but that condensation omits the reasons documents were cited. A methods paper can bridge topical groups for instrumental reasons. Lean too far toward compression and a map becomes an unwarranted claim of consensus; insist on full local context for every edge and the overview becomes hard to see. Diagnostic: Is a citation-structure map sufficient for the decision, or is the substantive purpose of particular citations essential?[1][3]
Structural–Framed Character¶
The portable skeleton is a binary relation whose membership is decided by intersection of two evidence neighborhoods. The named measure, however, is tied to scholarly documents and citation practice. Its evaluative weight is low at the definition level: the count records events without judging merit. Human-practice dependence is substantial because authors choose citations and citation conventions shape what is recorded. Its institutional origin lies in bibliometric indexing, although the set-intersection operation is mathematical. Its vocabulary travels into nearby citation-mapping methods, but “co-citation” outside document citation needs a newly declared unit, not just a metaphor. On import versus recognition, analysts import the counting rule into another bibliography and recognize its events there; they do not discover a citation-independent natural law in every association. Its character: structurally crisp as a relation and count, but framed by the domain's documentary and indexing practices; the generic relation travels farther than the named bibliometric measure.[1][3]
Structural Core vs. Domain Accent¶
The skeletal relation is a pairwise common-neighbor test: two focal nodes share an incoming neighbor under a directed link. That relation-level shape is why Relation is the proposed strict parent and Network a related graph lens. The domain-bound mechanism specifies what nodes and links mean: focal scholarly works, other works citing both, and a corpus-relative count of those citation events. Those bibliometric roles are not incidental ornament; remove them and one has a common-neighbor relation, not document co-citation. Therefore the named entry does not clear the prime bar, even though its formal skeleton can be recognized elsewhere.[1]
Instantiates / Related Primes¶
This entry is a kind of Relation. Joint citation defines a binary relation on pairs of cited works.
Relationships to Other Abstractions¶
Current abstraction Co-citation Domain-specific
Parents (1) — more general patterns this builds on
-
Co-citation is a kind of Relation Prime
Joint citation defines a binary relation on pairs of cited works.Two cited works stand in the co-citation relation if at least one work in a specified corpus cites both. The raw number of common citers weights the pair. This supplies the relata and inclusion rule of the live Relation prime while imposing a document-and-citation-specific predicate; it is not a claim that every relation is co-citation. Proposed only, pending independent DAG review.
Hierarchy path (1) — routes to 1 parentless root
- Co-citation → Relation
Neighborhood in Abstraction Space¶
Co-citation sits in a sparse region of the domain-specific corpus (67th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Faceted Classification & Metadata (14 abstractions)
Nearest neighbors
- Bibliographic Coupling — 0.85
- Citation Pointer — 0.85
- Yoked Control Design — 0.84
- Collostructional Analysis — 0.84
- Hendiadys — 0.83
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Bibliographic coupling asks whether a and b cite the same earlier work. Direct citation asks whether one of them cites the other. Textual similarity compares content without requiring a joint citer. Co-occurrence in a catalogue or keyword list likewise lacks the citing-document predicate. A co-citation cluster is an analytical grouping produced from multiple pair scores, not the score or relation for a single pair. The discriminating test is always: which documents in the declared corpus cite both focal works?[1][2]
References¶
[1] Henry Small, “Co-citation in the scientific literature: A new measure of the relationship between two documents”, Journal of the American Society for Information Science 24(4), 1973, pp. 265–269. Original publisher abstract and note 6 (citer-set intersection, raw and possible relative frequency) directly inspected; full article not assumed accessible. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w ↩x ↩y ↩z
[2] M. M. Kessler, “Bibliographic coupling between scientific papers”, American Documentation 14(1), 1963, pp. 10–25. Original publisher abstract inspected; outgoing-reference contrast is also explicitly made against the live bibliographic-coupling entry and Small's original differentiation. registry ↩a ↩b ↩c ↩d ↩e
[3] Henry Small and Belver C. Griffith, “The Structure of Scientific Literatures I: Identifying and Graphing Specialties”, Science Studies 4(1), 1974, pp. 17–40. Original publisher page, especially visible notes 17, 19 and 27, directly inspected; restricted full article not relied upon. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g