Skip to content

Neural scaling law

An empirical relation describing how neural-network loss or capability changes with model size, data, training compute or inference compute.

Version
v1 · 2026-09-08 · History
Domain-specific #
5758
Origin domain
machine learning
Subdomain
machine learning
Aliases
Scaling law for neural networks

Core Idea

Fits are regime- and architecture-dependent, extrapolation can fail at data quality or optimization transitions and benchmark contamination or selective reporting can mimic smooth scaling. Controlled training or evaluation runs vary one or jointly allocated resources, measured loss is fit to a power law or related curve and the fitted exponents predict marginal returns and compute-optimal allocation within a validated range. The abstraction is therefore identified by a declared carrier, a transformation or constraint over that carrier, and an invariant that tells an analyst whether the named structure is genuinely present.

Scope of Application

Neural scaling law belongs to machine learning and is useful where the analyst can specify the typed machine learning carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, then evaluate the model family and training recipe, parameter count data tokens training and inference compute, outcome metric and evaluation set, scaling variables held fixed or co-varied, fitted functional form exponents and uncertainty, observed range and extrapolation limit and compute-optimal allocation rule are explicit.

Clarity

The abstraction clarifies a crowded vocabulary by making the model family and training recipe, parameter count data tokens training and inference compute, outcome metric and evaluation set, scaling variables held fixed or co-varied, fitted functional form exponents and uncertainty, observed range and extrapolation limit and compute-optimal allocation rule are explicit the center of the account. A claim should name the carrier, the governing operation or relation, the applicable assumptions, and the recognition test.

Manages Complexity

Without the abstraction, an analyst must reason directly over many local details: the carrier roles, admissibility assumptions, competing conventions, derived invariants, boundary cases, and proof or validation obligations specific to Neural scaling law. Neural scaling law compresses them into the roles in the structural signature. That compression permits comparison across instances without erasing the variables that determine validity. It also exposes which details may be varied safely and which are constitutive.

Abstract Reasoning

  1. Identify the carrier. State what the elements, states, objects, or observations are: the typed machine learning carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets. Reject examples whose alleged carrier belongs to a different problem. 2. Lock the constitutive rule. Express the model family and training recipe, parameter count data tokens training and inference compute, outcome metric and evaluation set, scaling variables held fixed or co-varied, fitted functional form exponents and uncertainty, observed range and extrapolation limit and compute-optimal allocation rule are explicit independently of one notation or implementation.

Knowledge Transfer

Knowledge transfers strongly among subfields of machine learning because they reuse the typed machine learning carrier, including objects, relations, parameters, conventions, evidence, boundaries, and comparison targets, Controlled training or evaluation runs vary one or jointly allocated resources, measured loss is fit to a power law or related curve and the fitted exponents predict marginal returns and compute-optimal allocation within a validated range., and type the carrier, state every parameter and convention in the definition, test that the model family and training recipe, parameter count data tokens training and inference compute, outcome metric and evaluation set, scaling variables held fixed or co-varied, fitted functional form exponents and uncertainty, observed range and extrapolation limit and compute-optimal allocation rule are explicit, compare the nearest accepted identity, and report counterexamples, uncertainty, and limiting cases.

Relationships to Other Abstractions

Local relationship map for Neural scaling lawParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Neural scaling lawDOMAINPrime abstraction: Proportionality — is a kind ofProportionalityPRIME

Current abstraction Neural scaling law Domain-specific

Parents (1) — more general patterns this builds on

  • Neural scaling law is a kind of Proportionality Prime

    The proposed strict upward parent is prime:proportionality.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Neural scaling law sits in a crowded region of the domain-specific corpus (19th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Deep Learning Architectures & Scaling (16 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08