Ghost Character¶
A purportedly distinct written character entered in a character inventory or encoding through a source-lineage error rather than a securely attested character in the claimed tradition.
Core Idea¶
A ghost character is a purportedly distinct written character that acquired a place in a character inventory or encoding because a source was miscopied, misread, or wrongly identified, rather than because the claimed distinct character was securely attested in the original writing tradition. The error is about character identity and provenance, not merely whether a glyph can be displayed. Once an inventory gives that form a position, later dictionaries, mappings, fonts, and documents can make the form look well established even when their evidence all descends from the same mistaken listing.[1][2]
The diagnosis requires a source-lineage test. An unfamiliar character is not a ghost simply because an ordinary dictionary omits it; a rare form might be genuine. A later independent writing with the same shape does not necessarily validate the specific path by which an earlier list acquired it. Hiroyuki Sasahara's 1997 investigation analyzes ten forms absent from the original material cited during JIS selection and tests same-shaped forms in unrelated earlier writings; he does not claim that every questioned JIS form has the same origin. In a separate Tangut setting, a Unicode technical investigation calls U+1820D a ghost character while carefully qualifying that it has not been found in original Tangut sources as far as the investigators can ascertain.[1][3]
An encoding decision is not a historical authenticity certificate. The Unicode Standard itself acknowledges that a few erroneous Tangut dictionary characters are separately encoded, and the Tangut encoding proposal explains why inherited dictionary entries were not simply excluded on authenticity grounds. The entry can therefore remain technically usable even after its original character claim is questioned.[2][4]
Structural Signature¶
Sig role-phrases: purported distinct character identity → derivative inventory position → defective source lineage → source-critical comparison → possibly persistent downstream representation.
- Purported distinct character identity. A shape is represented as a separate character in the relevant writing system, not just as damaged ink, a font rendering glitch, or an acknowledged shape variant. Without this identity claim, the case is a typographic accident but not yet a ghost character.[1][2]
- Derivative character inventory. A source list, dictionary head entry, or character repertoire gives the shape a stable place that subsequent users can cite or encode. The inventory turns a local error into a reproducible identity claim. Japanese JIS selection and modern Tangut lexicographic entries are unlike institutional routes to this role.[1][2]
- Source-lineage defect. Comparison with the purported original materials exposes a copying, reading, or attribution path that fails to support the claimed distinct character. A truly independent, correctly identified use in the claimed lineage would defeat a confirmed ghost diagnosis. Mere absence from one familiar book does not establish the defect.[1][3]
- Source-critical comparison. The reviewer follows the form backward through intermediate lists to primary evidence and distinguishes visually identical forms reached through independent paths. This investigation recognizes the ghost; it is not necessarily part of the historical event that introduced it.[1]
- Stable downstream representation. A standard position or code point can persist because existing text and mappings depend on it. Persistence is a common consequence, not a requirement that every ghost be copied forever or eventually deleted. Sasahara describes compatibility constraints in the JIS setting; Unicode acknowledges separate encoding of erroneous Tangut forms.[1][4]
The five roles should not be compressed into “a character nobody uses.” Current use of a code point and historical source authenticity answer different questions. Conversely, uncertainty about attestation warrants suspected or unresolved status, not automatic conviction.
What It Is Not¶
It is not a rare character as such. A local place-name graph or manuscript form can be sparsely documented and still genuine. Sasahara's 1997 paper establishes a narrower result: ten selected forms were absent from the claimed originals, yet matching shapes could be found in unrelated earlier materials. Those independent matches do not repair the broken JIS source path, and neither his ten-form result nor a modern dictionary omission licenses calling every disputed JIS item a ghost.[1]
It is not simply a glyph variant or a font defect. Variant shapes can be legitimately unified or separately represented according to a repertoire's source rules. A bad display of an otherwise correct character is a rendering problem. Ghost status concerns a purported distinct character identity that the claimed source chain fails to justify. The Tangut proposal explicitly treats variant unification separately from erroneous dictionary forms.[2]
It is not a ghost word. A ghost word is a spurious lexical word-form; a ghost character is a character-inventory identity error. A word may comprise several characters, and a character need not have a stand-alone lexical reading. The frozen seed's attempt to put dictionary words, encoded characters and phantom islands under “Ghost Word” crosses this boundary by analogy rather than by the literal named definition. The earlier hold is retained in the bundle as provenance, not promoted as another accepted identity.[5]
It is not proof that Unicode or JIS should delete the code point. Historical criticism and encoding maintenance have different objectives. Preserving old data can justify retaining a code point while marking its provenance as erroneous or disputed.[1][2][4]
Scope of Application¶
The literal scope is curation of written-character repertoires and the source documents from which they are built: archival or administrative character lists, dictionaries, character standards, and encoded script inventories. The abstraction applies when a derivative item is presented as a separate character and a provenance investigation can test whether the purported original evidence actually supports it.[1][2]
It does not depend on one script or one standards body. Sasahara studies Japanese JIS X 0208 selection against earlier source lists; the Tangut standards proposal and Unicode Technical Note examine modern Tangut dictionary entries against original Tangut sources. The common object is a character identity moving through documentary chains, even though the error mechanisms and institutions differ.[1][3]
The boundary is epistemic as well as technical. A source-traced copy error can be classified strongly. A form merely not yet located in surviving originals must retain the investigators' qualification; for U+1820D, West and Zaytsev say “as far as we can ascertain,” not that every lost Tangut document has been searched. Conversely, a demonstrable later use of an encoded form shows modern circulation, not necessarily pre-encoding authenticity.[3]
Clarity¶
The ghost-character distinction untangles three claims that otherwise collapse: the shape is representable, the shape appears in a derivative reference, and the purported distinct character existed in the alleged original source. A code point proves the first; a dictionary head entry proves the second; neither alone proves the third. A provenance chain, not merely a font sample or search-result count, tests the identity claim.[1][4]
It also clarifies why a “matching” old glyph may be misleading. Sasahara's study specifically asks whether earlier same-shaped forms in unrelated documents are genuine transmission evidence or independent coincidences of jitai (character form). Visual agreement can coexist with different origin and meaning. Without the source-lineage distinction, one accidental match would seem to settle the character's provenance when it does not.[1]
Manages Complexity¶
An inherited character inventory can involve source lists, editors, copied dictionaries, glyph drawings, code positions and later user documents. The abstraction compresses this mass into a small diagnostic: what identity is claimed, where did the entry first acquire authority, and does the cited primary source support that exact identity? That reduction lets an investigator separate source error from rarity and from later technical persistence.[1][2]
The reduction does not replace philology. It identifies which links must be checked. In the Tangut case the proposal describes misread or miswritten forms replicated among modern dictionaries; the later technical note isolates one code point and compares modern dictionary history to original Tangut attestation. Treating all entries as one “bad Unicode character” would obscure the distinct evidence for each.[2][3]
Abstract Reasoning¶
Start with the putative identity and the first inventory position that treats it as distinct. Follow the cited source backward rather than counting derivative repetitions as independent attestations. Compare the specific glyph and its use in the alleged original documents, allowing for genuine rare forms, independently convergent shapes and transcription error. Then label the result supported, error-derived, or unresolved rather than forcing a yes/no decision on incomplete evidence.[1]
This procedure licenses two separate conclusions. The historical conclusion concerns whether the cited lineage established a distinct character. The engineering conclusion concerns how to preserve existing coded text. Those can differ: an error-derived item may remain encoded for compatibility. A downstream example of the code point is therefore evidence of present representation, not by itself evidence of the original character claim.[1][2][4]
Knowledge Transfer¶
The test transfers literally between Japanese JIS kanji-source investigation and Tangut repertoire construction: both require a supposed character, an authoritative derivative listing, an asserted historical source, and a comparison that can distinguish error from genuine independent use. Their scripts, documents and standards processes differ; the same role relation remains.[1][2][3]
The wider pattern of a false reference entry acquiring authority also resembles ghost words and phantom cartographic objects, but that is analogy at a higher level, not literal membership in the Ghost Character class. Live Traceability offers a related provenance infrastructure and Standard a related maintained artifact, yet neither has the required ghost-character source-error signature. A generic erroneous-inventory-entry abstraction is a future-prime question, not a reason to rename this character-specific identity or infer a strict DAG parent.
Examples¶
Japanese JIS source investigation. Sasahara follows questioned forms in the original JIS selection back through intermediate data to the source lists claimed for them; he specifically examines forms absent from the cited originals and asks whether earlier same-shaped forms in other documents are independent coincidences rather than actual source transmission. Purported distinct character identity: the questioned form is treated as its own kanji in the JIS repertoire. Derivative character inventory: its JIS row-cell position was produced from compiled character lists. Source-lineage defect: for the error-derived subset, the intermediary form is not supported by the purported original list. Source-critical comparison: the investigator compares original materials, derivative lists and unrelated same-shaped uses; this also protects genuinely rare characters from false ghost labeling. Stable downstream representation: established JIS positions are hard to replace without affecting compatibility. Mapped back: the diagnostic is a broken claimed path from source text to inventory identity, not merely that the graph looked strange to a modern reader. This is a class-level case from the original investigator, not an unverified claim that every often-listed JIS “ghost” shares one exact mistake.[1][6]
Tangut U+1820D. West and Zaytsev identify this encoded ideograph as a ghost, saying no original Tangut source attests it as far as they can ascertain; they trace its presence in modern lexicography to discussion by earlier scholars. Purported distinct character identity: U+1820D is separately identified as a Tangut ideograph. Derivative character inventory: a modern Tangut dictionary lists it and Unicode assigns it a position. Source-lineage defect: the investigation finds no original Tangut text supporting that purported distinct identity, while later scholarly transmission explains the modern entry. Source-critical comparison: the technical note compares modern dictionary material and earlier commentary with original-source attestation, preserving its stated limit. Stable downstream representation: the character remains separately encoded despite the ghost judgment. Mapped back: an apparently authoritative coded entry and a historical character attestation are distinct; the former persists even when the latter is unestablished in the investigators' search.[3][4]
Structural Tensions¶
Historical authenticity versus encoding compatibility. Removing a demonstrated erroneous item could make a repertoire more faithful to primary script evidence, but established code positions and mapped documents impose a cost on deletion or substitution. Retaining it preserves access to old encoded data while making the mistake look more authoritative to readers who equate encoding with legitimacy. Diagnostic: Is the decision about classification of origin or about continued representation of existing data? Make the two decisions separately.[1][2][4]
Skeptical source testing versus recovery of rare characters. A strict “not in my dictionary” filter is cheap but rejects real low-frequency place-name or historical forms. Accepting every derivative list entry is also cheap but carries errors forward. Primary-source tracing takes more work and can still leave uncertainty, yet it is what distinguishes rare, error-derived and merely unresolved forms. Diagnostic: Does the source presented for this exact character survive inspection, or are all hits derivative copies or independent same-shape coincidences?[1][3]
Structural–Framed Character¶
Ghost Character lies toward the framed end of the structural–framed spectrum, with a real transferable diagnostic inside character-repertoire curation. Evaluative weight: “ghost” is a judgment about the warrant for a character identity, not a neutral description of glyph rarity; it should be qualified when source evidence is incomplete. Human-practice dependence: scholars, editors and standards committees decide which forms count as separate entries and document their provenance, so the category depends on those practices. Institutional origin: the term's clearest cases arise from JIS/Unicode and lexicographic inventories, though no single institution owns the error pattern. Vocabulary travel: the phrase can name erroneous entries in distinct scripts, but crossing to words or islands needs a new genus, not silent literal transfer. Import versus recognition: once an independently supported source chain is compared with a derivative character listing, an investigator can recognize the same error form in a new script; importing the label on mere unfamiliarity would be irresponsible.[1][2][4]
Its character: a domain-specific, source-critical class of writing-system identity anomalies. Its structure travels across character inventories, while its diagnosis remains framed by the script's evidence, the inventory's identity rules and an explicitly qualified historical judgment.
Structural Core vs. Domain Accent¶
Core: a derivative authoritative list presents an item as distinct, while tracing it to claimed originals reveals that the particular identity came from error rather than the asserted source. Domain accent: the item is a written character, whose glyph shape, historical attestation, dictionary treatment and code-point status are separate objects of analysis. Removing those character-specific distinctions yields a general erroneous-entry story but no longer the named Ghost Character class.[1][2]
The portable source-error and authority-propagation skeleton is an explicit future-prime question. Live Traceability could support backward investigation but is not a strict parent: a ghost can arise before anyone maintains bidirectional provenance records. Standardization can entrench a position, but the ghost is not itself the act of convergence on a norm. This domain-specific identity does not clear the prime bar by borrowing their breadth.
Instantiates / Related Primes¶
No parent prime is asserted. Traceability is related because a sufficiently documented origin chain helps test ghost status; it is not necessary to the creation of the erroneous character. Encoding And Decoding is related to coded characters but its paired content-to-code and code-to-content process does not define an error-derived character identity. Standardization explains some persistence and coordination effects after an inventory is adopted, not the ghost-character class itself.
The proposed DAG therefore leaves this identity unparented. It may later connect to a carefully defined higher-order erroneous-entry or source-lineage abstraction if that parent is itself admitted and its necessary-genus relation is justified.
Neighborhood in Abstraction Space¶
Ghost Character sits in a sparse region of the domain-specific corpus (65th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Language, Mind & Meaning-Making (57 abstractions)
Nearest neighbors
- Lexical analysis — 0.85
- The Chosen One (Trope) — 0.85
- Databending — 0.84
- GI (complexity) — 0.84
- Writing system — 0.84
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Do not equate this with Standard, the normative artifact containing a repertoire; with Phantom Inventory, a physical-stock mismatch; or with Phantom reference, a Java reference type. The shared words “ghost” and “phantom” are lexical, not graph edges. Do not alias the Japanese-only expression ghost kanji to this broader cross-script node without a vocabulary review; it is a narrower member term and remains a proposal in the candidate ledger.
References¶
[1] Hiroyuki Sasahara, “字体に生じる偶然の一致:「JIS X0208」と他文献における字体の『暗合』と『衝突』” [“Coincidence and Clash of Jitai”], 日本語科学 1 (1997): 7–24, esp. §3 and conclusion, DOI 10.15084/00001963. Original investigator paper in the National Institute for Japanese Language and Linguistics repository. https://repository.ninjal.ac.jp/record/1979/files/kk_ngkgk_001_01.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v
[2] Andrew West et al., “Proposal to Encode the Tangut Script in the UCS,” JTC1/SC2/WG2 N4522 (2014), §5.2.1, “Glyph Variants,” paragraphs on erroneous dictionary forms and encoding policy. https://www.unicode.org/wg2/docs/n4522.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n
[3] Andrew West and Viacheslav Zaytsev, “Tangut Character Additions and Glyph Corrections,” Unicode Technical Note #42, version 2 (2019), §3.1.2 on U+1820D and p. 49 source table. This technical note is scholarly guidance, not normative Unicode text. https://www.unicode.org/notes/tn42/tn42-2.pdf registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h
[4] Unicode Consortium, The Unicode Standard, version 16.0, Chapter 18, Tangut “Encoding Principles,” including erroneous or “ghost” dictionary characters separately encoded. https://unicode.org/versions/Unicode16.0.0/core-spec/chapter-18/ registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h
[5] Walter W. Skeat, “Report upon Ghost-Words, or Words Which Have No Real Existence,” Transactions of the Philological Society (1885–87), original address pp. 351–354. Used only for the separate lexical boundary, not as a source for encoded characters. https://archive.org/stream/transact188500philuoft/transact188500philuoft_djvu.txt registry ↩
[6] National Institute for Japanese Language and Linguistics, original repository record for Sasahara (1997), with author, publication and DOI metadata. https://repository.ninjal.ac.jp/records/1979 registry ↩