Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.¶
Northcutt, C. G., Athalye, A., & Mueller, J. (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. Proceedings of NeurIPS Datasets and Benchmarks Track.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Ground Truth
- Suppose inter-annotator studies show the human labels themselves are only 96% correct (a measurable noise floor).
This sourceQuantifies label errors in benchmark test sets (≈3.4% average; ≈6% in the ImageNet validation set), validated by multi-rater human review — a measured reference noise floor.
- Suppose inter-annotator studies show the human labels themselves are only 96% correct (a measurable noise floor).
- Reference Standard Decay
- In ML evaluation, the apparatus is a model-scoring pipeline; the role-bearing reference is a benchmark's gold-standard label set (e.g., a held-out test set with "correct" answers); models report accuracy or F1 against that reference.
This sourceShows gold-standard benchmark label sets contain errors and re-labeling shifts the reference against which models are scored.
- In ML evaluation, the apparatus is a model-scoring pipeline; the role-bearing reference is a benchmark's gold-standard label set (e.g., a held-out test set with "correct" answers); models report accuracy or F1 against that reference.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 1 other way.
Registry ID ref:7f10d292016e · see in the full table