Skip to content

Label noise

Incorrect, inconsistent, ambiguous, or corrupted target labels in supervised-learning data, arising randomly or systematically from annotators, processes, proxies, attacks, or changing definitions.

Version
v1 · 2026-09-08 · History
Domain-specific #
5241
Origin domain
machine learning data quality
Subdomain
machine learning data quality

Core Idea

Label noise changes empirical risk and can drive memorization, miscalibration, subgroup error, and biased evaluation; instance-independent, class-conditional, and instance-dependent noise require different identification assumptions. A latent or adjudicated target passes through a labeling channel whose error distribution can depend on true class, features, annotator, time, or adversary, producing the observed training label. The abstraction is therefore identified by a declared carrier, a transformation or constraint over that carrier, and an invariant that tells an analyst whether the named structure is genuinely present.

Scope of Application

Label noise belongs to machine learning data quality and is useful where the analyst can specify the typed machine learning data quality carrier, defining objects and relations, parameters, conventions, evidence, boundary cases, and comparison targets, then evaluate the task and label ontology, latent-truth assumption, annotation process, observed labels, noise taxonomy and transition model, repeated or gold labels, class and subgroup prevalence, train-test contamination, detection method, uncertainty, and correction evaluation are explicit. The scope is broad within that domain but bounded by the need for the task and label ontology, latent-truth assumption, annotation process, observed labels, noise taxonomy and transition model, repeated or gold labels, class and subgroup prevalence, train-test contamination, detection method, uncertainty, and correction evaluation are explicit.

Clarity

The abstraction clarifies a crowded vocabulary by making the task and label ontology, latent-truth assumption, annotation process, observed labels, noise taxonomy and transition model, repeated or gold labels, class and subgroup prevalence, train-test contamination, detection method, uncertainty, and correction evaluation are explicit the center of the account. A claim should name the carrier, the governing operation or relation, the applicable assumptions, and the recognition test.

Manages Complexity

Without the abstraction, an analyst must reason directly over many local details: the carrier roles, admissibility assumptions, competing conventions, derived invariants, boundary cases, and proof or validation obligations specific to Label noise. Label noise compresses them into the roles in the structural signature. That compression permits comparison across instances without erasing the variables that determine validity. It also exposes which details may be varied safely and which are constitutive.

Abstract Reasoning

  1. Identify the carrier. State what the elements, states, objects, or observations are: the typed machine learning data quality carrier, defining objects and relations, parameters, conventions, evidence, boundary cases, and comparison targets. Reject examples whose alleged carrier belongs to a different problem. 2.

Knowledge Transfer

Knowledge transfers strongly among subfields of machine learning data quality because they reuse the typed machine learning data quality carrier, defining objects and relations, parameters, conventions, evidence, boundary cases, and comparison targets, A latent or adjudicated target passes through a labeling channel whose error distribution can depend on true class, features, annotator, time, or adversary, producing the observed training label., and type the carrier, state every parameter and convention in the definition, test that the task and label ontology, latent-truth assumption, annotation process, observed labels, noise taxonomy and transition model, repeated or gold labels, class and subgroup prevalence, train-test contamination, detection method, uncertainty, and correction evaluation are explicit, compare the nearest accepted identity, and report counterexamples, uncertainty, and limiting cases.

Relationships to Other Abstractions

Local relationship map for Label noiseParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Label noiseDOMAINPrime abstraction: Data Integrity — is a kind ofData IntegrityPRIME

Current abstraction Label noise Domain-specific

Parents (1) — more general patterns this builds on

  • Label noise is a kind of Data Integrity Prime

    The proposed strict upward parent is prime:data_integrity.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Label noise sits in a crowded region of the domain-specific corpus (30th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Machine Learning & Statistical Estimation (24 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08