AI and the Everything in the Whole Wide World Benchmark.¶
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021). AI and the Everything in the Whole Wide World Benchmark. Proceedings of the NeurIPS Datasets and Benchmarks Track.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Construct Validity
- The classical validity battery ports almost unchanged from psychometrics to the evaluation of computational systems: convergent validity becomes "does the benchmark agree with held-out human judgment?", discriminant becomes "is it uncorrelated with confounds like input length?"
This sourceArgues that 'general' AI benchmarks have construct-validity problems, including failure to discriminate the named capability from confounds, and calls for checking benchmark validity rather than assuming it.
- The classical validity battery ports almost unchanged from psychometrics to the evaluation of computational systems: convergent validity becomes "does the benchmark agree with held-out human judgment?", discriminant becomes "is it uncorrelated with confounds like input length?"
- Fallacy Of Misplaced Concreteness
- Statistics and ML — benchmark scores or held-out accuracy mistaken for the underlying capability they were meant to measure.
This sourceArgues benchmark scores are mistaken for the underlying capability they were meant to measure — construct reification in machine learning.
- Statistics and ML — benchmark scores or held-out accuracy mistaken for the underlying capability they were meant to measure.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 1 other way.
Registry ID ref:3c1f83d878a0 · see in the full table