LLM Evaluators Recognize and Favor Their Own Generations.¶
Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems (NeurIPS).
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Measurement
- The unit and calibration chain maps onto the requirement that the benchmark be anchored to an independent reference (human-judged answers); a model used as an automated judge without such an anchor reproduces the benchmark's blind spots, which is the metrology failure of calibrating an instrument against itself.
This sourceShows an automated LLM judge exhibits self-preference/self-recognition bias—evidence that an evaluator calibrated against itself reproduces its own blind spots, the metrology failure of calibrating an instrument against itself.
- The unit and calibration chain maps onto the requirement that the benchmark be anchored to an independent reference (human-judged answers); a model used as an automated judge without such an anchor reproduces the benchmark's blind spots, which is the metrology failure of calibrating an instrument against itself.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:745d6c51852e · see in the full table