Fisher Kernel¶
A model-based similarity equal to the Fisher-information-normalized inner product of two observations' log-likelihood gradients under a common fitted generative model.
Core Idea¶
The Fisher kernel represents an observation by how it would change a generative model. For model p(X|theta), the Fisher score is the gradient of log likelihood with respect to theta at a shared reference estimate. Two observations point in similar parameter directions when they produce aligned score vectors.
Scope of Application¶
- Sequence classification. Hidden Markov or related models turn variable-length sequence likelihood gradients into fixed kernel comparisons.
- Document retrieval. Probabilistic text models supply score representations for discriminative similarity or ranking.
- Image representation. Local descriptors can be pooled through a mixture model into Fisher-style explicit features.
- Hybrid learning. A generative model contributes structured representation while a kernel machine supplies a discriminative boundary.
Clarity¶
A reproducible construction names p(X|theta), the training data and estimate of theta, parameter coordinates, score computation, Fisher-information estimate, regularization, and any explicit-feature approximation. Distance from a kernel value is a downstream choice. Reparameterization claims depend on using the metric correctly rather than dropping the information normalization.
Manages Complexity¶
The kernel compresses a structured observation into sensitivity across model parameters. This converts different lengths and internal alignments into a fixed comparison space while retaining which latent mechanisms each example stresses. The compression inherits misspecification: distinctions ignored by the generative model cannot be recovered simply by a powerful classifier.
Abstract Reasoning¶
- Choose a differentiable generative model appropriate to the observations' structure.
- Fit or otherwise fix one common reference parameter vector.
- Compute each observation's gradient of log likelihood with respect to those parameters.
- Estimate and regularize the Fisher information, documenting approximations.
- Form the weighted score inner product or a mathematically equivalent feature map.
Knowledge Transfer¶
The method transfers across data types when a shared differentiable likelihood and valid information metric exist. Using gradients from an arbitrary loss may form a useful tangent feature but is not automatically a Fisher kernel. The general transfer pattern is representing examples by their pressure on a fitted model.
Relationships to Other Abstractions¶
Current abstraction Fisher Kernel Domain-specific
Parents (1) — more general patterns this builds on
-
Fisher Kernel is a kind of Positive-definite kernel Domain-specific
The Fisher Kernel is a Positive-Definite Kernel built from Fisher-information-normalized likelihood-score feature vectors.
Hierarchy path (1) — routes to 1 parentless root
- Fisher Kernel → Positive-definite kernel → Function (Mapping)
Neighborhood in Abstraction Space¶
Fisher Kernel sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Statistical Hypothesis Tests & Diagnostics (9 abstractions)
Nearest neighbors
- Approximate Bayesian Computation — 0.89
- Neural modeling fields — 0.88
- Autoencoder — 0.85
- Interval Predictor Model — 0.84
- Misuse of p-values — 0.84
Computed from structural-signature embeddings · 2026-10-08