Training Compute-Optimal Large Language Models¶
Hoffmann, J., Borgeaud, S., & Mensch, A. (2022). Training Compute-Optimal Large Language Models. Advances in Neural Information Processing Systems 35.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Diminishing Incremental Gains
- Liebig's Law of the Minimum
- In AI training, model performance caps on whichever of data quantity, data quality, compute, parameters, or inference budget is binding at the current scale, and the compute-optimal literature is essentially a search for the Liebig configuration.
This sourceThe Chinchilla result showing prior large models were parameter-rich but data-starved, a Liebig diagnosis over training inputs.
- In AI training, model performance caps on whichever of data quantity, data quality, compute, parameters, or inference budget is binding at the current scale, and the compute-optimal literature is essentially a search for the Liebig configuration.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:a4d07aab7655 · see in the full table