Garbage in, garbage out? Do machine learning application papers in social computing report where human-labeled training data comes from?¶
Geiger, R. S., Yu, K., Yang, Y., Dai, M., Qiu, J., Tang, R., & Huang, J. (2020). Garbage in, garbage out? Do machine learning application papers in social computing report where human-labeled training data comes from?. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT, 325-336.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Native-Category Flattening
- The asymmetry of recoverability bites hard and expensively — a model trained on these labels can never learn the mid-purchase-versus-overnight distinction, because it was destroyed before training; recovering it requires re-annotating from raw text, if the raw text was even retained.
This sourceDocuments how fixed annotation label sets applied to data carrying finer native distinctions constrain every downstream model and resist after-the-fact recovery.
- The asymmetry of recoverability bites hard and expensively — a model trained on these labels can never learn the mid-purchase-versus-overnight distinction, because it was destroyed before training; recovering it requires re-annotating from raw text, if the raw text was even retained.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:dd0c006c9237 · see in the full table