Investigating Data Contamination in Modern Benchmarks for Large Language Models.¶
Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 8706-8719.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Data Leakage
- And in adversarial settings, published test benchmarks eventually leak into the training data of the systems they were meant to evaluate.
This sourceDocuments that published evaluation benchmarks leak into the pretraining corpora of the LLMs they were meant to evaluate, inflating reported scores — the adversarial benchmark-contamination case of data leakage.
- And in adversarial settings, published test benchmarks eventually leak into the training data of the systems they were meant to evaluate.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:76fbe383f4fb · see in the full table