Machine Learning Models & Representations¶
← Back to Domain-Specific Families
Abstractions about how machine-learning systems learn, represent, and generalize from data, spanning neural-network architecture components (hidden layer, attention, Neural Turing machine), learning paradigms (transfer learning, zero-shot learning, ensemble learning), and theoretical or evaluative constructs (concept class, neural scaling law, out-of-bag error).
25 abstractions in this family — domain-specific abstractions that sit near one another in structural-signature space (k-means over structural-signature embeddings). Each is shown with its short description.
- Attention (machine learning) — A neural-network mechanism that computes context-dependent weights over representations and combines them so each output can focus selectively on relevant inputs.
- Automatic basis function construction — Learning task-reusable basis functions that compress a large state space for value-function approximation.
- Bag-of-words model in computer vision — An image representation that quantizes local visual descriptors into a learned vocabulary and summarizes the image by a histogram of visual-word occurrences.
- Binary classification — Assign observations to exactly two declared classes through a learned or specified decision rule, keeping scores, thresholds, reference labels, asymmetric errors, prevalence, and evaluation population distinct.
- Concept class — A family of Boolean-valued concepts or classifiers over a common instance domain, serving as the hypothesis universe in computational learning theory.
- Decision tree pruning — The removal or replacement of low-value branches from a decision tree to reduce complexity and improve expected generalization.
- Deep learning — Learn task-relevant hierarchical representations with multilayer parameterized neural networks trained end to end by optimization over data, enabling complex prediction and generation at the cost of opacity, data dependence, and compute.
- Ensemble learning — A machine-learning strategy combining predictions from multiple models so their complementary errors yield a stronger aggregate predictor.
- Error-driven learning — Learning that updates expectations or parameters in proportion to a discrepancy between predicted and observed outcomes, so surprising events produce larger representational change.
- Hidden layer — A neural-network layer situated between inputs and outputs whose learned nonlinear transformations construct intermediate representations used by later layers.
- Hybrid Kohonen self-organizing map — A neural architecture coupling a self-organizing map front end to supervised hidden and output layers.
- Kernel principal component analysis — Nonlinear dimensionality reduction obtained by performing PCA in an implicit reproducing-kernel feature space.
- Label noise — Incorrect, inconsistent, ambiguous, or corrupted target labels in supervised-learning data, arising randomly or systematically from annotators, processes, proxies, attacks, or changing definitions.
- Large width limits of neural networks — Asymptotic regimes in which neural-network layer widths tend to infinity and random networks converge to analytically tractable Gaussian-process, kernel, mean-field, or feature-learning descriptions.
- Lazy learning — A machine-learning strategy that postpones generalization from stored training examples until a prediction query arrives.
- Linear separability — The property that two labeled point sets lie on opposite sides of at least one affine hyperplane.
- Logistic model tree — Partition predictor space with a decision tree while fitting logistic-regression models through the tree, yielding piecewise probabilistic classification whose local linear logits are induced, inherited, and pruned together.
- Multiple instance learning — A supervised-learning setting in which labels attach to bags of instances while instance-level labels are absent or only indirectly constrained.
- Neural scaling law — An empirical relation describing how neural-network loss or capability changes with model size, data, training compute or inference compute.
- Neural Turing machine — A differentiable recurrent architecture coupling a neural controller to an addressable external memory.
- Out-of-bag error — A predictive-error estimate computed for each training case using only bagged models that excluded it from their bootstrap samples.
- Proto-value function — A task-independent spectral basis function learned from a state-transition graph to approximate value functions in reinforcement learning.
- Swish function — A smooth neural-network activation family fβ(x)=x·sigmoid(βx) that interpolates between a scaled linear map and a ReLU-like gate while remaining mildly nonmonotonic for positive β.
- Transfer learning — A machine-learning strategy that reuses representations, parameters or examples learned in a source task or domain to improve learning in a related target task.
- Zero-shot learning — A learning setup that predicts classes absent from training by transferring through auxiliary semantic descriptions or attributes shared with seen classes.