Stabilizing Transformer Training by Preventing Attention Entropy Collapse¶
Zhai, S., Likhomanenko, T., & Littwin, E. (2023). Stabilizing Transformer Training by Preventing Attention Entropy Collapse.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Emphasis
- Attention mechanisms in machine learning exhibit this pathology as "attention collapse," requiring regularization and normalization to prevent uniform or degenerate distributions.
This sourceDocuments attention entropy collapse — attention distributions becoming degenerate/peaked — and shows regularization/normalization mitigations; supports the machine-learning 'attention collapse' half of T3 (emphasis inflation / collapse to no-emphasis).
- Attention mechanisms in machine learning exhibit this pathology as "attention collapse," requiring regularization and normalization to prevent uniform or degenerate distributions.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:d90d5f2c6696 · see in the full table