Infomax¶
An information-theoretic design principle that selects an admissible input–output mapping by maximizing their mutual information under a stated probability model.
Core Idea¶
Infomax chooses or learns an input–output mapping by the mutual information retained between a modeled input and its output. It requires an input distribution, admissible mappings, an output relation and an information-based selection criterion. A fixed MI measurement is not itself a design problem, and a particular gradient rule is not part of the identity.[^ref-1a2f2c54542a]
Scope of Application¶
Linsker formulated the principle for allowed linear transformations under noise and resource limits. Bell and Sejnowski used a nonlinear information-maximizing network for blind separation of mixed signals. Deep InfoMax used related input–representation objectives for image encoders, including local/global choices that affect downstream usefulness. These share the selection criterion, not an identical algorithm or guaranteed independent output.[ref-1a2f2c54542a][ref-cf70a331afdf][^ref-72bb983005c2]
Clarity¶
State the input ensemble, mapping family and noise, sampling or resolution model before asking which map maximizes MI. A continuous deterministic, zero-noise formulation can require careful entropy treatment; unbounded mappings can make a claimed optimum ill-posed. “Retains more information” also does not mean “best for every later task.”[ref-1a2f2c54542a][ref-cf70a331afdf][^ref-27105170891b]
Manages Complexity¶
The criterion compares unlike filters and encoders through one information-based choice structure. The comparison remains honest only when its model-specific constraints and estimated objective are visible: a contrastive or multi-view score may be a qualified surrogate, not the same quantity automatically.[ref-1a2f2c54542a][ref-27105170891b]
Abstract Reasoning¶
Choose \(f\) from an allowed family \(\mathcal F\) to increase or maximize \(I(X;Y_f)\) under a stated joint model. This is a specialized optimization problem: \(f\) is the choice, mutual information is the objective and \(\mathcal F\) defines feasibility. Exact attainment, training procedure, independent components and downstream utility are separate questions.[ref-1a2f2c54542a][ref-cf70a331afdf]
Knowledge Transfer¶
Blind audio separation and image representation learning retain the same input–mapping–output–information roles while using different signals and estimators. General best-choice reasoning is already live Optimization; Infomax remains domain-specific because Shannon-information and probabilistic representation assumptions give the named pattern its distinctive meaning.[ref-cf70a331afdf][ref-72bb983005c2]
[^ref-1a2f2c54542a]: Ralph Linsker, “An Application of the Principle of Maximum Information Preservation to Linear Systems”, Advances in Neural Information Processing Systems 1 (1988), PDF pp. 1–3. [^ref-cf70a331afdf]: Anthony J. Bell and Terrence J. Sejnowski, “An Information-Maximization Approach to Blind Separation and Blind Deconvolution”, Neural Computation 7(6) (1995), 1129–1159, PDF pp. 1–3. [^ref-72bb983005c2]: R Devon Hjelm et al., “Learning Deep Representations by Mutual Information Estimation and Maximization”, ICLR (2019), arXiv:1808.06670v5, PDF pp. 1–4. [^ref-27105170891b]: Michael Tschannen et al., “On Mutual Information Maximization for Representation Learning”, ICLR (2020), arXiv:1907.13625, §§1–4.
Relationships to Other Abstractions¶
Current abstraction Infomax Domain-specific
Parents (1) — more general patterns this builds on
-
Infomax is a kind of Optimization Prime
Infomax selects admissible mappings by an input–output mutual-information objective.
Hierarchy path (1) — routes to 1 parentless root
- Infomax → Optimization
Neighborhood in Abstraction Space¶
Infomax sits in a sparse region of the domain-specific corpus (66th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Statistical Learning & Model Failure Modes (41 abstractions)
Nearest neighbors
- Signal Quantization — 0.86
- Machine-Learning Model — 0.84
- Physical-System Model — 0.84
- Learnable Function Class — 0.84
- Kriging — 0.83
Computed from structural-signature embeddings · 2026-10-08