Skip to content

Cache Language Model

A language model that combines a general next-word distribution with a distribution learned from recent text history to adapt its predictions as the text unfolds.

Version
v1 · 2026-10-07 · History
Domain-specific #
13822
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomains
Natural Language Processing, Statistical Language Modeling → Computer Science & Software Engineering
Aliases
Cache Based Language Model

Core Idea

A cache language model adapts next-word predictions to the text it has just processed. It combines a general probability model trained earlier with a local probability estimate from recent words or their contexts. A repeated word may receive more weight when the local history supports it, while the general model still covers words absent from that history. The term refers to prediction, not just faster computer-memory access.[ref-5c3e5c9191b3][ref-28b055b0f737]

Scope of Application

In speech recognition, the recent words of a dictated document can supply a local signal in addition to a fixed language model. Jelinek and colleagues tested such a cache with IBM's TANGORA recognizer on 14 dictated office documents from five speakers. They updated the cache with corrected words after each utterance and cleared it between documents. The reported recognition gains belong to that limited test.[^ref-5c3e5c9191b3]

In neural language modeling, Grave and colleagues stored earlier hidden-state/next-word pairs. The model compared the current context with those stored contexts and combined their word probabilities with a pre-trained recurrent model. Their reported improvement was lower perplexity on language-model datasets, not a measured translation or speech-recognition error reduction.[^ref-28b055b0f737]

Clarity

Ask what is stored, how it changes word probabilities, and what outcome was measured. The IBM cache used recent word counts; the neural cache used similarity to stored contexts. Both produced a local next-word distribution and combined it with a broad baseline. Storing text without changing prediction is merely a buffer. A processor cache that speeds computation is a different use of “cache.”[ref-5c3e5c9191b3][ref-28b055b0f737]

Manages Complexity

The model separates adaptation into four choices: a general language model, a recent-history cache, a way to estimate local probabilities, and a rule for combining the two. This lets a system adapt to one document without retraining its broad model for every new topic. Window length, update rules and mixing weight still matter. The TANGORA experiment used corrected text; a cache filled with recognition mistakes would need separate evaluation.[ref-5c3e5c9191b3][ref-28b055b0f737]

Abstract Reasoning

For a candidate next word, compare the general model's probability with the evidence in the current cache. If the candidate appears in recent or similar contexts, the local estimate may shift the combined probability toward it. If it is absent, the general model still contributes. Presence in the cache does not guarantee that every word's final probability rises. Check the task and history quality before increasing the cache's influence.[ref-5c3e5c9191b3][ref-28b055b0f737]

Knowledge Transfer

The literal idea transfers from count-based n-gram models to neural models: both use a history-derived local distribution beside a general distribution. The count estimator and hidden-state similarity method do not transfer as the same implementation. The live Conditional Probability Prime supplies the broader idea of predicting a word given prior text; this entry adds the language-specific cache architecture and has a strict composition/presupposes edge to that Prime.[ref-5c3e5c9191b3][ref-28b055b0f737]

Example

Jelinek and colleagues combined a fixed trigram model with a cache of recent word frequencies. In their TANGORA recognition test, corrected dictated text updated the cache after each utterance. Mapped back: the fixed trigram is the general distribution, recent document words form the cache, smoothed word counts give the local estimate, and interpolation combines them. Their observed error-rate reduction does not prove a universal benefit.[^ref-5c3e5c9191b3]

Grave and colleagues instead combined a pre-trained recurrent model with a cache of hidden states and following words. Similarity between a current and earlier state supplied the local word estimate. Mapped back: the recurrent model is the baseline, stored state/word pairs are the cache, similarity weights form the local distribution, and interpolation or global normalization combines the sources. Their tested outcome was language-model perplexity.[^ref-28b055b0f737]

Relationships to Other Abstractions

Local relationship map for Cache Language ModelParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Cache Language ModelDOMAINPrime abstraction: Conditional Probability — presupposesConditionalProbabilityPRIME

Current abstraction Cache Language Model Domain-specific

Parents (1) — more general patterns this builds on

  • Cache Language Model presupposes Conditional Probability Prime

    Its general, recent-history, and combined predictions are probabilities of a next word conditional on prior text; without that relation the model cannot perform its defining adaptation.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Cache Language Model sits in a sparse region of the domain-specific corpus (99th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

Fast computation caching stores results for reuse without necessarily changing word probabilities. Offline retraining changes a model using another corpus without a live recent-history distribution. Factored Language Model represents tokens through declared linguistic factors and a parent graph; it is a different model form. Guaranteed rare-word correction is too strong: the cited studies report specific results under specific update and evaluation conditions.[ref-5c3e5c9191b3][ref-28b055b0f737]

References

[^ref-5c3e5c9191b3]: F. Jelinek, B. Merialdo, S. Roukos and M. Strauss, “A Dynamic Language Model for Speech Recognition,” Speech and Natural Language: Proceedings of a Workshop Held at Pacific Grove, California, February 19–22, 1991 (1991), PDF pp. 1–3, especially “Cache Language Model,” “Isolated Speech Recognition,” and Tables 1–4. https://aclanthology.org/H91-1057.pdf

[^ref-28b055b0f737]: Edouard Grave, Armand Joulin and Nicolas Usunier, “Improving Neural Language Models with a Continuous Cache,” arXiv:1612.04426 (2016), §§2–3, 5 and Tables 1–2. https://arxiv.org/html/1612.04426