Ghost Character¶
A purportedly distinct written character entered in a character inventory or encoding through a source-lineage error rather than a securely attested character in the claimed tradition.
Core Idea¶
A ghost character is a purportedly distinct written character that entered a character inventory or encoding through a source-lineage error, not a securely attested distinct character in the source tradition claimed for it. The mistake may begin in copying, reading or attributing a source list; the inventory can then give the form an apparently authoritative position. The test is provenance, not whether modern software can display the glyph or whether an ordinary dictionary lists it.[ref-b6dd6172def7][ref-db4bb81b4fcd]
An unfamiliar form is only suspect until original-source review distinguishes error from a genuine rare character. Sasahara's 1997 paper examines ten selected JIS forms absent from the cited original materials; independent earlier uses of matching shapes do not necessarily establish that original JIS lineage. His result does not classify every questioned JIS form. In a different script, a Unicode investigation calls Tangut U+1820D a ghost while carefully saying that no original Tangut attestation was found as far as the investigators could ascertain.[ref-b6dd6172def7][ref-e94e7917d432]
Scope of Application¶
The literal setting is written-character repertoire curation: source lists, dictionaries, standards and encoded script inventories. In Japanese JIS X 0208 investigation, Sasahara compared derivative character lists with alleged original materials and separated copying errors from independent same-shaped uses. In Tangut, modern dictionary forms were evaluated against original script sources; the encoding proposal acknowledged that some ghost characters had been copied among dictionaries and still treated those head entries as encoding candidates.[ref-b6dd6172def7][ref-db4bb81b4fcd]
Thus a code point can survive a historical ghost diagnosis. Unicode's Tangut chapter says that a few erroneous or ghost dictionary characters are separately encoded. Compatibility and representability are engineering questions; original character authenticity is a historical one. This entry does not claim that every obscure character is erroneous, that every ghost must be deleted, or that a later use of an encoded form proves its original historical lineage.[ref-976ebe7c0cbf][ref-b6dd6172def7]
Clarity¶
The concept separates a displayable glyph, a derivative reference entry, and an independently supported character identity. The first two can be verified by seeing a font or dictionary. Only a source-chain investigation tests the third. A visually identical mark in an unrelated old source may be an accidental convergence of shape rather than evidence that the questioned list copied a genuine character from that source.[^ref-b6dd6172def7]
It also distinguishes a ghost character from a ghost word: a spurious lexical word-form is not a kind of encoded character. The frozen seed conflated the two and added phantom islands by analogy; this focused reroute covers only the originally selected Ghost characters candidate. The broader phantom-entry analogy and the separate Ghost Word identity remain unpromoted proposals.[^ref-ed927344c3af]
Manages Complexity¶
Instead of counting every later dictionary, font and web use as fresh evidence, trace the asserted character back to the first source path that gave it an independent identity. Ask which intermediary copied which original, what the original actually shows, and whether later matches are dependent copies or independent uses. This compresses a large documentary trail into an auditable identity judgment without pretending that all doubtful entries have one origin story.[ref-b6dd6172def7][ref-db4bb81b4fcd]
The Japanese and Tangut settings share this diagnostic but not necessarily one exact error mechanism. They also show why the historical diagnosis should be kept separate from maintenance of old encoded text: deleting an established position can make existing mappings unusable even when the original entry was a mistake.[ref-b6dd6172def7][ref-976ebe7c0cbf]
Abstract Reasoning¶
Identify the putative distinct character and its inventory position; follow its cited provenance through intermediate lists to the claimed original; inspect the original form and context; then classify the case as supported, error-derived, or unresolved. A rare legitimate character should remain supported even if familiar dictionaries omit it. A derivative chain should not be counted as multiple independent attestations. Qualified conclusions such as the Tangut U+1820D finding should keep their evidentiary limit.[ref-b6dd6172def7][ref-e94e7917d432]
For a confirmed ghost, decide separately whether the entry should be documented, annotated, retained for compatibility or treated differently in a future repertoire. Historical source criticism does not automatically prescribe a particular encoding change.[ref-db4bb81b4fcd][ref-976ebe7c0cbf]
Knowledge Transfer¶
The abstraction transfers literally across Japanese JIS kanji and Tangut Unicode repertoire work: both have a claimed distinct character, a derivative authoritative listing, and a test against original-source attestation. Sasahara's JIS investigation and the Tangut U+1820D analysis fill those same roles with unrelated scripts and documents.[ref-b6dd6172def7][ref-e94e7917d432]
Traceability is a related infrastructure for following origins, not a strict parent of the ghost-character class; Standard is an artifact that may contain such an entry, not its genus. The more general pattern of an error gaining reference authority is a future-prime question. Applying “ghost character” to a word or phantom island would be analogy, not literal transfer.
[^ref-b6dd6172def7]: Hiroyuki Sasahara, “字体に生じる偶然の一致:「JIS X0208」と他文献における字体の『暗合』と『衝突』” [“Coincidence and Clash of Jitai”], 日本語科学 1 (1997): 7–24, esp. §3 and conclusion, DOI 10.15084/00001963. https://repository.ninjal.ac.jp/record/1979/files/kk_ngkgk_001_01.pdf [^ref-db4bb81b4fcd]: Andrew West et al., “Proposal to Encode the Tangut Script in the UCS,” JTC1/SC2/WG2 N4522 (2014), §5.2.1. https://www.unicode.org/wg2/docs/n4522.pdf [^ref-e94e7917d432]: Andrew West and Viacheslav Zaytsev, “Tangut Character Additions and Glyph Corrections,” Unicode Technical Note #42, version 2 (2019), §3.1.2 on U+1820D. https://www.unicode.org/notes/tn42/tn42-2.pdf [^ref-976ebe7c0cbf]: Unicode Consortium, The Unicode Standard, version 16.0, Chapter 18, Tangut “Encoding Principles.” https://unicode.org/versions/Unicode16.0.0/core-spec/chapter-18/ [^ref-ed927344c3af]: Walter W. Skeat, “Report upon Ghost-Words, or Words Which Have No Real Existence,” Transactions of the Philological Society (1885–87), pp. 351–354. This supports only the lexical boundary, not character provenance. https://archive.org/stream/transact188500philuoft/transact188500philuoft_djvu.txt
Neighborhood in Abstraction Space¶
Ghost Character sits in a sparse region of the domain-specific corpus (65th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Language, Mind & Meaning-Making (57 abstractions)
Nearest neighbors
- Lexical analysis — 0.85
- The Chosen One (Trope) — 0.85
- Databending — 0.84
- GI (complexity) — 0.84
- Writing system — 0.84
Computed from structural-signature embeddings · 2026-10-08