Benchmark Data Contamination of Large Language Models¶
Xu, C., Guan, S., Greene, D., & Kechadi, M. (2024). Benchmark Data Contamination of Large Language Models: A Survey.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Construct Validity
- The real leak is the construct-to-proxy bridge: discriminant validity fails because pass-rate correlates strongly with whether near-identical problems appeared in training data (contamination), so the benchmark is partly measuring memorization, a distinct construct.
This sourceSurveys how train-test contamination inflates measured benchmark capability via memorization rather than genuine capability — a distinct construct.
- The real leak is the construct-to-proxy bridge: discriminant validity fails because pass-rate correlates strongly with whether near-identical problems appeared in training data (contamination), so the benchmark is partly measuring memorization, a distinct construct.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:a4d4394a0598 · see in the full table