Assessing agreement on classification tasks¶
Carletta, J. (1996). Assessing agreement on classification tasks: the kappa statistic.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Domain-specific¶
- Inter-Annotator Agreement
- Computational linguistics and ML dataset construction — the home turf, where IAA gates whether a labelled corpus (sentiment, NER spans, discourse relations, image classes) is released or its scheme revised
This sourceCarletta's argument that computational linguistics should adopt the content-analysis kappa statistic as its reliability measure for classification tasks. Carletta's case that raw agreement is uninterpretable without chance correction, with kappa's .67/.8 content-analysis conventions for acceptable reliability.
Supported in partVerified against a saved copy of the source
“if researchers can’t even show that different people can agree about the judgments on which their research is based, then there is no chance of replicating the research results.”
- Computational linguistics and ML dataset construction — the home turf, where IAA gates whether a labelled corpus (sentiment, NER spans, discourse relations, image classes) is released or its scheme revised
Verification¶
Does it exist? Not checked yet. This entry carries no identifier to resolve. It was extracted from the citation as written in the article, normalized, and deduplicated against the rest of the registry.
Does it back the claim? Read against the text for 1 of 1 citation: 1 supported in part. Each verdict is shown under its citation below, with what in the work backs the sentence.
Support is checked per citation rather than per work — the same source can be cited soundly in one article and wrongly in another. Per-citation recording began recently, so a citation with no recorded check is a gap in the record rather than evidence it went unchecked.
See how references were verified.
Registry ID ref:96bd4be78fbd · see in the full table