Cache Language Model¶
A language model that combines a general next-word distribution with a distribution learned from recent text history to adapt its predictions as the text unfolds.
Core Idea¶
A cache language model adapts next-word predictions to the text it has just processed. It combines a general probability model trained earlier with a local probability estimate from recent words or their contexts. A repeated word may receive more weight when the local history supports it, while the general model still covers words absent from that history. The term refers to prediction, not just faster computer-memory access.[ref-5c3e5c9191b3][ref-28b055b0f737]
Scope of Application¶
In speech recognition, the recent words of a dictated document can supply a local signal in addition to a fixed language model. Jelinek and colleagues tested such a cache with IBM's TANGORA recognizer on 14 dictated office documents from five speakers. They updated the cache with corrected words after each utterance and cleared it between documents. The reported recognition gains belong to that limited test.[^ref-5c3e5c9191b3]
In neural language modeling, Grave and colleagues stored earlier hidden-state/next-word pairs. The model compared the current context with those stored contexts and combined their word probabilities with a pre-trained recurrent model. Their reported improvement was lower perplexity on language-model datasets, not a measured translation or speech-recognition error reduction.[^ref-28b055b0f737]
Clarity¶
Ask what is stored, how it changes word probabilities, and what outcome was measured. The IBM cache used recent word counts; the neural cache used similarity to stored contexts. Both produced a local next-word distribution and combined it with a broad baseline. Storing text without changing prediction is merely a buffer. A processor cache that speeds computation is a different use of “cache.”[ref-5c3e5c9191b3][ref-28b055b0f737]
Manages Complexity¶
The model separates adaptation into four choices: a general language model, a recent-history cache, a way to estimate local probabilities, and a rule for combining the two. This lets a system adapt to one document without retraining its broad model for every new topic. Window length, update rules and mixing weight still matter. The TANGORA experiment used corrected text; a cache filled with recognition mistakes would need separate evaluation.[ref-5c3e5c9191b3][ref-28b055b0f737]
Abstract Reasoning¶
For a candidate next word, compare the general model's probability with the evidence in the current cache. If the candidate appears in recent or similar contexts, the local estimate may shift the combined probability toward it. If it is absent, the general model still contributes. Presence in the cache does not guarantee that every word's final probability rises. Check the task and history quality before increasing the cache's influence.[ref-5c3e5c9191b3][ref-28b055b0f737]
Knowledge Transfer¶
The literal idea transfers from count-based n-gram models to neural models: both use a history-derived local distribution beside a general distribution. The count estimator and hidden-state similarity method do not transfer as the same implementation. The live Conditional Probability Prime supplies the broader idea of predicting a word given prior text; this entry adds the language-specific cache architecture and has a strict composition/presupposes edge to that Prime.[ref-5c3e5c9191b3][ref-28b055b0f737]
Example¶
Jelinek and colleagues combined a fixed trigram model with a cache of recent word frequencies. In their TANGORA recognition test, corrected dictated text updated the cache after each utterance. Mapped back: the fixed trigram is the general distribution, recent document words form the cache, smoothed word counts give the local estimate, and interpolation combines them. Their observed error-rate reduction does not prove a universal benefit.[^ref-5c3e5c9191b3]
Grave and colleagues instead combined a pre-trained recurrent model with a cache of hidden states and following words. Similarity between a current and earlier state supplied the local word estimate. Mapped back: the recurrent model is the baseline, stored state/word pairs are the cache, similarity weights form the local distribution, and interpolation or global normalization combines the sources. Their tested outcome was language-model perplexity.[^ref-28b055b0f737]
Relationships to Other Abstractions¶
Current abstraction Cache Language Model Domain-specific
Parents (1) — more general patterns this builds on
-
Cache Language Model presupposes Conditional Probability Prime
Its general, recent-history, and combined predictions are probabilities of a next word conditional on prior text; without that relation the model cannot perform its defining adaptation.
Hierarchy paths (2) — routes to 2 parentless roots
- Cache Language Model → Conditional Probability → Probability → Measure → Aggregation → Micro Macro Linkage
- Cache Language Model → Conditional Probability → Probability → Measure → Set and Membership
Neighborhood in Abstraction Space¶
Cache Language Model sits in a sparse region of the domain-specific corpus (99th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Unclustered & Miscellaneous (2551 abstractions)
Nearest neighbors
- Factored Language Model — 0.77
- Kneser–Ney Smoothing — 0.75
- Heaps' Law — 0.75
- Postings List — 0.74
- Structural Risk Minimization — 0.74
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Fast computation caching stores results for reuse without necessarily changing word probabilities. Offline retraining changes a model using another corpus without a live recent-history distribution. Factored Language Model represents tokens through declared linguistic factors and a parent graph; it is a different model form. Guaranteed rare-word correction is too strong: the cited studies report specific results under specific update and evaluation conditions.[ref-5c3e5c9191b3][ref-28b055b0f737]
References¶
[^ref-5c3e5c9191b3]: F. Jelinek, B. Merialdo, S. Roukos and M. Strauss, “A Dynamic Language Model for Speech Recognition,” Speech and Natural Language: Proceedings of a Workshop Held at Pacific Grove, California, February 19–22, 1991 (1991), PDF pp. 1–3, especially “Cache Language Model,” “Isolated Speech Recognition,” and Tables 1–4. https://aclanthology.org/H91-1057.pdf
[^ref-28b055b0f737]: Edouard Grave, Armand Joulin and Nicolas Usunier, “Improving Neural Language Models with a Continuous Cache,” arXiv:1612.04426 (2016), §§2–3, 5 and Tables 1–2. https://arxiv.org/html/1612.04426