Skip to content

Self-supervised learning

Training a model with supervisory targets derived from the data itself.

Version
v1 · 2026-09-28 · History
Domain-specific #
11956
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomains
Machine Learning, Self Supervised Representation Learning → Computer Science & Software Engineering
Aliases
SSL, Self-supervised representation learning

Core Idea

Self-supervised learning creates a training signal from the data being learned. A model may compare two transformed views of one item or reconstruct a hidden part of it; in either case the target relation is generated without an external annotator for that pretraining objective. Optimization updates model parameters, ideally yielding representations useful beyond the pretext task.

SimCLR's image-view contrast is a defining construction, while masked autoencoders provide a separate published image-reconstruction use. Both papers report downstream labeled evaluations. That later use of class labels is compatible with a self-supervised pretraining stage and must be separated from the target source during training. The family is broader than contrastive learning alone, and merely possessing unlabeled data without a derived objective is not an instance.

Scope of Application

The self-generated-target claim applies to the pretraining stage, not necessarily to later evaluation.

  • Computer vision. Learn visual features from large image collections.
  • Speech and language. Predict masked or future signal parts.
  • Representation learning. Pretrain reusable embeddings before task-specific supervision.
  • Benchmark evaluation. Separate pretraining objective from downstream label use.

Clarity

Self-supervised learning builds training targets from its own data. SimCLR uses paired image views; a masked autoencoder reconstructs hidden patches. Both update model parameters without human class labels for pretraining, yet later labeled evaluation or fine-tuning remains possible. Augmentation alone does not make a human-labeled classifier self-supervised.

Manages Complexity

Target generation replaces manual annotation with an internal relation, but the target design controls what the model can learn. Augmentations may remove useful details; masked reconstruction may focus on pixel statistics. Researchers therefore evaluate representations on separate tasks and must not treat success on a pretext loss as automatic transfer.

Abstract Reasoning

Name the stage and data, generate a target from each input, optimize model parameters against it, then evaluate later tasks while keeping any label use stage-specific.

Knowledge Transfer

The data-derived-target pattern travels across images, text, audio and other modalities, but a literal SSL instance requires an updating model and a target generated without external labels for that stage. A domain-specific proxy task may change while these roles remain.

Relationships to Other Abstractions

Local relationship map for Self-supervised learningParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Self-supervisedlearningDOMAINPrime abstraction: Learning — is a kind ofLearningPRIME

Current abstraction Self-supervised learning Domain-specific

Parents (1) — more general patterns this builds on

  • Self-supervised learning is a kind of Learning Prime

    Data-derived targets drive durable model-parameter updates that alter later predictions.

Hierarchy paths (2) — routes to 2 parentless roots

Neighborhood in Abstraction Space

Self-supervised learning sits in a crowded region of the domain-specific corpus (33rd percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Communication, Learning & Information Practices (15 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08