Skip to content

Machine Learning Models & Representations

← Back to Domain-Specific Families

Abstractions about how machine-learning systems learn, represent, and generalize from data, spanning neural-network architecture components (hidden layer, attention, Neural Turing machine), learning paradigms (transfer learning, zero-shot learning, ensemble learning), and theoretical or evaluative constructs (concept class, neural scaling law, out-of-bag error).

25 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.

  • Attention (machine learning) — A neural-network mechanism that computes context-dependent weights over representations and combines them so each output can focus selectively on relevant inputs.
  • Automatic basis function construction — Learning task-reusable basis functions that compress a large state space for value-function approximation.
  • Bag-of-words model in computer vision — An image representation that quantizes local visual descriptors into a learned vocabulary and summarizes the image by a histogram of visual-word occurrences.
  • Binary classification — Assign observations to exactly two declared classes through a learned or specified decision rule, keeping scores, thresholds, reference labels, asymmetric errors, prevalence, and evaluation population distinct.
  • Concept class — A family of Boolean-valued concepts or classifiers over a common instance domain, serving as the hypothesis universe in computational learning theory.
  • Decision tree pruning — The removal or replacement of low-value branches from a decision tree to reduce complexity and improve expected generalization.
  • Deep learning — Learn task-relevant hierarchical representations with multilayer parameterized neural networks trained end to end by optimization over data, enabling complex prediction and generation at the cost of opacity, data dependence, and compute.
  • Ensemble learning — A machine-learning strategy combining predictions from multiple models so their complementary errors yield a stronger aggregate predictor.
  • Error-driven learning — Learning that updates expectations or parameters in proportion to a discrepancy between predicted and observed outcomes, so surprising events produce larger representational change.
  • Hidden layer — A neural-network layer situated between inputs and outputs whose learned nonlinear transformations construct intermediate representations used by later layers.
  • Hybrid Kohonen self-organizing map — A neural architecture coupling a self-organizing map front end to supervised hidden and output layers.
  • Kernel principal component analysis — Nonlinear dimensionality reduction obtained by performing PCA in an implicit reproducing-kernel feature space.
  • Label noise — Incorrect, inconsistent, ambiguous, or corrupted target labels in supervised-learning data, arising randomly or systematically from annotators, processes, proxies, attacks, or changing definitions.
  • Large width limits of neural networks — Asymptotic regimes in which neural-network layer widths tend to infinity and random networks converge to analytically tractable Gaussian-process, kernel, mean-field, or feature-learning descriptions.
  • Lazy learning — A machine-learning strategy that postpones generalization from stored training examples until a prediction query arrives.
  • Linear separability — The property that two labeled point sets lie on opposite sides of at least one affine hyperplane.
  • Logistic model tree — Partition predictor space with a decision tree while fitting logistic-regression models through the tree, yielding piecewise probabilistic classification whose local linear logits are induced, inherited, and pruned together.
  • Multiple instance learning — A supervised-learning setting in which labels attach to bags of instances while instance-level labels are absent or only indirectly constrained.
  • Neural scaling law — An empirical relation describing how neural-network loss or capability changes with model size, data, training compute or inference compute.
  • Neural Turing machine — A differentiable recurrent architecture coupling a neural controller to an addressable external memory.
  • Out-of-bag error — A predictive-error estimate computed for each training case using only bagged models that excluded it from their bootstrap samples.
  • Proto-value function — A task-independent spectral basis function learned from a state-transition graph to approximate value functions in reinforcement learning.
  • Swish function — A smooth neural-network activation family fβ(x)=x·sigmoid(βx) that interpolates between a scaled linear map and a ReLU-like gate while remaining mildly nonmonotonic for positive β.
  • Transfer learning — A machine-learning strategy that reuses representations, parameters or examples learned in a source task or domain to improve learning in a related target task.
  • Zero-shot learning — A learning setup that predicts classes absent from training by transferring through auxiliary semantic descriptions or attributes shared with seen classes.