Towards Monosemanticity: Decomposing Language Models with Dictionary Learning¶
Bricken, T., Templeton, A., & Batson, J. (2023). Towards Monosemanticity: Decomposing Language Models with Dictionary Learning: Decomposing Language Models with Dictionary Learning. Transformer Circuits Thread.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Sparse Coding
- In machine learning, an L1 or KL sparsity penalty on hidden activations produces units with specific, interpretable triggers, and sparse autoencoders are now used to recover monosemantic features from the dense activations of transformer residual streams.
This sourceUses sparse autoencoders with an L1 penalty to recover human-readable monosemantic features from the dense, polysemantic activations of a transformer residual stream.
- In machine learning, an L1 or KL sparsity penalty on hidden activations produces units with specific, interpretable triggers, and sparse autoencoders are now used to recover monosemantic features from the dense activations of transformer residual streams.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:3ac620666163 · see in the full table