Discovering Language Model Behaviors with Model-Written Evaluations¶
Perez, E., Ringer, S., & Lukošiūtė, K. (2022). Discovering Language Model Behaviors with Model-Written Evaluations.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Unreliable Narrator
- Machine-generated text — fluent output is systematically shaped by training distribution, preference tuning, prompt sycophancy, and confabulation; the right protocol maintains an explicit distortion model rather than treating output as ground truth.
This sourceDemonstrates that RLHF-tuned model outputs are systematically shaped by preference tuning — notably sycophancy, where larger models increasingly mirror the user's stated view — establishing the distortion model machine output requires.
- Machine-generated text — fluent output is systematically shaped by training distribution, preference tuning, prompt sycophancy, and confabulation; the right protocol maintains an explicit distortion model rather than treating output as ground truth.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:ea4de3515fbf · see in the full table