Scaling Laws for Neural Language Models.¶
Kaplan, J., McCandlish, S., Henighan, T., & Brown, T. B. (2020). Scaling Laws for Neural Language Models.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Aggregate-Marginal Divergence
- AI compute and conservation. Aggregate model performance improves while marginal benefit per training unit collapses along the scaling curve, and protected-area acreage grows while each added hectare is more remote and less biodiverse.
This sourceDocuments diminishing marginal returns to compute/data/parameters along power-law scaling curves while aggregate capability keeps rising — the forced aggregate-marginal divergence.
- AI compute and conservation. Aggregate model performance improves while marginal benefit per training unit collapses along the scaling curve, and protected-area acreage grows while each added hectare is more remote and less biodiverse.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:e63f05d88926 · see in the full table