Scaling Laws & Growth Patterns¶
← Back to Domain-Specific Families
Abstractions that describe regular scaling patterns and frequency statistics in data and language, spanning empirical scaling laws (Benford's, Heaps', and Gibrat's laws), growth notation (big O, L-notation), lexical frequency effects (type-token ratio, word frequency effect), and count-based sequence models (Kneser-Ney smoothing, prediction by partial matching).
12 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.
- Benford's Law — Score a dataset's honesty by checking whether its leading digits follow the fixed logarithmic curve log₁₀((d+1)/d) — about 30% start with 1, only 5% with 9 — that scale-spanning multiplicative data must obey.
- Big O Notation — Classify a function by its order of growth — discarding constant factors and small-input detail — so algorithms can be compared by how they scale rather than how fast they run on one machine.
- Chow–Liu Tree — Approximate a discrete joint distribution by the tree-factorized model whose edges maximize total pairwise mutual information.
- Factored Language Model — A statistical language model that represents each token as a bundle of linguistic factors and predicts a selected factor from a configurable graph of prior lexical, morphological, syntactic, or semantic factors, using generalized backoff when contexts are sparse.
- Gibrat's Law — The claim that a firm's proportional growth rate is independent of its current size — which, iterated as multiplicative iid noise, makes log-size a random walk and drives the cross-sectional size distribution toward log-normal.
- Heaps' Law — The empirical regularity that a corpus's distinct-word-type count grows as a sublinear power of the token count, V = K·N^β with β below 1, so vocabulary rises without bound ever more slowly and can be extrapolated across scale from two fitted parameters.
- Kneser–Ney Smoothing — An n-gram probability-estimation method that discounts observed counts and backs off using how widely a word continues distinct contexts, not just its total frequency.
- L-notation — A two-parameter asymptotic scale that records the subexponential exponent and leading logarithmic constant of number-theoretic algorithms.
- Ninety-Ninety Rule — The first 90 percent of the code takes the first 90 percent of the time and the last 10 percent takes the other 90 percent — a deadpan warning that nominal-progress metrics measure only the estimable bulk-work population and never the non-parallelizable, heavy-tailed completion tail.
- Prediction by Partial Matching — Estimate the next symbol from adaptive continuation statistics at the longest informative recent context, backing off to shorter contexts when the continuation is unseen.
- Type-Token Ratio — Estimate a language sample's lexical diversity as distinct word types over total tokens, while correcting for the sample-size confound that makes vocabulary grow sublinearly and raw TTR fall with length.
- Word Frequency Effect — The robust finding that high-frequency words are recognised faster and more accurately than low-frequency ones, because lexical access is a graded, exposure-keyed threshold — retrieval speed falling roughly with the logarithm of a word's lifetime corpus frequency for that reader.