Skip to content

Fisher Kernel

A model-based similarity equal to the Fisher-information-normalized inner product of two observations' log-likelihood gradients under a common fitted generative model.

Version
v1 · 2026-09-28 · History
Domain-specific #
9475
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomains
Statistical Machine Learning, Kernel Methods → Computer Science & Software Engineering
Aliases
Fisher score kernel, Fisher information kernel

Core Idea

The Fisher kernel represents an observation by how it would change a generative model. For model p(X|theta), the Fisher score is the gradient of log likelihood with respect to theta at a shared reference estimate. Two observations point in similar parameter directions when they produce aligned score vectors.

Scope of Application

  • Sequence classification. Hidden Markov or related models turn variable-length sequence likelihood gradients into fixed kernel comparisons.
  • Document retrieval. Probabilistic text models supply score representations for discriminative similarity or ranking.
  • Image representation. Local descriptors can be pooled through a mixture model into Fisher-style explicit features.
  • Hybrid learning. A generative model contributes structured representation while a kernel machine supplies a discriminative boundary.

Clarity

A reproducible construction names p(X|theta), the training data and estimate of theta, parameter coordinates, score computation, Fisher-information estimate, regularization, and any explicit-feature approximation. Distance from a kernel value is a downstream choice. Reparameterization claims depend on using the metric correctly rather than dropping the information normalization.

Manages Complexity

The kernel compresses a structured observation into sensitivity across model parameters. This converts different lengths and internal alignments into a fixed comparison space while retaining which latent mechanisms each example stresses. The compression inherits misspecification: distinctions ignored by the generative model cannot be recovered simply by a powerful classifier.

Abstract Reasoning

  1. Choose a differentiable generative model appropriate to the observations' structure.
  2. Fit or otherwise fix one common reference parameter vector.
  3. Compute each observation's gradient of log likelihood with respect to those parameters.
  4. Estimate and regularize the Fisher information, documenting approximations.
  5. Form the weighted score inner product or a mathematically equivalent feature map.

Knowledge Transfer

The method transfers across data types when a shared differentiable likelihood and valid information metric exist. Using gradients from an arbitrary loss may form a useful tangent feature but is not automatically a Fisher kernel. The general transfer pattern is representing examples by their pressure on a fitted model.

Relationships to Other Abstractions

Local relationship map for Fisher KernelParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Fisher KernelDOMAINDomain-specific abstraction: Positive-definite kernel — is a kind ofPositive-defini…DOMAIN

Current abstraction Fisher Kernel Domain-specific

Parents (1) — more general patterns this builds on

  • Fisher Kernel is a kind of Positive-definite kernel Domain-specific

    The Fisher Kernel is a Positive-Definite Kernel built from Fisher-information-normalized likelihood-score feature vectors.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Fisher Kernel sits in a moderately populated region (55th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Statistical Hypothesis Tests & Diagnostics (9 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08