Skip to content

Cramér–Rao Estimator Efficiency

Compare an unbiased scalar estimator's variance with its regular-model Cramér–Rao information bound, under explicit conditions.

Version
v1 · 2026-10-03 · History
Domain-specific #
13104
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Estimation Theory, Parametric Inference → Experimental Design & Statistics

Core Idea

Cramér–Rao estimator efficiency is a finite-sample, model-relative comparison for an unbiased scalar estimator. In a regular parametric model, Fisher information \(I_n(\theta)\) for the entire sample gives the variance lower bound \(1/I_n(\theta)\). If estimator \(T\) of \(\theta\) is unbiased and has positive finite variance, define its efficiency at \(\theta\) by

\[e_\theta(T)=\frac{1/I_n(\theta)}{\operatorname{Var}_\theta(T)}.\]

Under the bound's conditions and positive finite information, \(0<e_\theta(T)\le1\). Equality at a given \(\theta\) says that \(T\) attains the information bound there. A claim of finite-sample efficiency for a model class often demands equality for every parameter value, which is stronger than pointwise equality. The bound might not be attainable by any estimator in a particular class, so the ratio measures distance from a mathematical lower bound, not necessarily removable variance slack relative to an actually feasible rival.[1]

The reusable abstraction is typed estimator and model → information-derived lower bound → actual sampling variance → conditional ratio. Normal-location and Bernoulli-proportion models instantiate it with different data and likelihoods. The broader frozen article “Efficiency (statistics)” also treats relative efficiency, asymptotic criteria, tests and designs. The broad name Efficiency is already the live prime and is not used as this domain-specific slug.[1][2][3]

Structural Signature

Sig role-phrases:

  • Regular scalar model and target parameter: fixes likelihood, sample size and \(\theta\); the information benchmark depends on this model. The simple inequality needs support and differentiation conditions.
  • Unbiased estimator and sample: supplies the data-to-estimate rule \(T(X_1,\ldots,X_n)\) being rated within the same experiment. Biased estimators need a different comparison or modified theorem.
  • Information variance lower bound: supplies \(1/I_n(\theta)\), a theorem-qualified benchmark—not a guarantee some estimator attains it.
  • Actual estimator variance: measures the spread of this rule at this parameter under this model. It must not be swapped for MSE without accounting for bias.
  • Bound-to-variance ratio: reports the pointwise score; a whole-model verdict requires checking all \(\theta\) in the stated parameter space.[1][2]

If model support changes with \(\theta\), the standard score/information equalities can fail and the simple bound may be invalid. If the estimator and information refer to different sample sizes or parameters, the ratio loses its meaning. If \(T\) is biased, small variance is not enough to infer small squared error. These are constitutive checks, not optional footnotes.[1][2]

What It Is Not

It is not generic statistical efficiency across tests, designs and estimators; those objects have different targets and benchmarks. It is not relative efficiency of two procedures, which compares their variances or asymptotic sample-size needs under a declared numerator convention and can exceed one. It is not asymptotic efficiency, which asks about limiting variance as sample size grows. The frozen normal mean-versus-median \(2/\pi\) figure is an asymptotic relative comparison, not this score for an arbitrary finite \(n\).[2][3]

An estimator with minimum variance among unbiased estimators does not automatically attain the Cramér–Rao bound; a lower bound may be loose. Conversely, attainment is stronger than being merely one favorable comparison. Do not interpret \(e=1\) at one parameter value as a uniform claim. Do not call an estimator “inefficient” in all senses from this ratio alone: it may be biased with lower MSE, more robust under model misspecification, or cheaper to compute.[1][2]

Live Efficiency asks whether an actual feasible alternative dominates an input–output arrangement under constraints. This statistical ratio shares benchmark language, but an unattainable information bound is not necessarily a feasible alternative or a Pareto frontier. Lexical overlap does not license a strict prime parent here.

Scope of Application

The displayed ratio applies to a scalar target parameter, an unbiased estimator with positive finite variance, a positive finite Fisher information, and the regularity needed for the Rao–Cramér inequality—such as parameter-independent support and justified interchange of differentiation and expectation in the lecture's formulation. For IID observations, \(I_n=nI_1\) under those assumptions. With a vector parameter, nuisance parameters, bias or irregular support, the benchmark and matrix comparison must be rederived; this scalar score should not simply be copied.[1]

For normal location with known variance, the sample mean attains the bound. For Bernoulli success probability $0<p<1\(, the sample proportion does likewise. Both are exact finite-sample cases in different statistical carriers. MIT's Uniform\)[0,\theta]$ example is a decisive exclusion test: its support depends on \(\theta\), the regular information equality fails, and an unbiased maximum-based estimator can beat the naively computed reciprocal-information expression. That does not refute the regular theorem; it shows why its hypotheses matter.[1]

An estimator's efficiency can depend on \(\theta\) and on the assumed model. A claim at one point, across an entire parameter set, or asymptotically as \(n\to\infty\) must be labeled accordingly. The score does not certify reliability if the likelihood model is wrong.[1][2]

Clarity

This abstraction separates three statements often compressed into “best estimator”: (1) a lower bound is valid under regularity, (2) a particular estimator reaches it, and (3) the estimator is best under the decision's actual loss and model uncertainty. The first is a theorem about a class, the second an equality calculation, and the third can require bias, robustness or other criteria. Neither (1) nor (2) entails (3) without added assumptions.[1][2]

It also makes the denominator visible. If two unbiased rules use the same \(n\) observations, the same model and the same target, their scores can be compared by their variances. If one uses a different parameter, model or number of observations, its information bound changes too. The numerical label alone is not a portable ranking.

Manages Complexity

The score compresses a whole estimator-sampling distribution into two quantities: an information lower bound and an actual variance. In a regular model, the ratio gives a quick dimensionless statement about precision relative to the bound. For the two worked models, it immediately distinguishes using all \(n\) observations from discarding \(n-1\) of them: $1$ versus \(1/n\) under the same experimental frame.[1]

The compression omits bias, tail risk, computational cost and robustness. It also omits whether the bound is attainable in the estimator class. Thus a score far below one identifies a gap to the theorem's lower bound, but not necessarily a constructive way to eliminate that gap. Preserve those limitations whenever the score informs estimator choice.[2]

Abstract Reasoning

Specify the parametric model, target \(\theta\), observed sample and estimator \(T\). Check unbiasedness and positive finite variance. Derive the score and information \(I_n(\theta)\) from the same likelihood, including the support and differentiation assumptions. Prove the Cramér–Rao inequality applies; then calculate \(e_\theta(T)\). If it equals one, state whether that is pointwise or throughout \(\Theta\). If it is below one, do not infer the missing precision is attainable without a separate construction or existence result.[1]

As a boundary check, repeat the first step for a model whose support depends on \(\theta\). MIT's Uniform\([0,\theta]\) calculation shows that formally differentiating a likelihood while ignoring changing support can manufacture an invalid information benchmark. The disciplined inference is to stop and rederive an appropriate bound, not to report a ratio greater than one as a paradox.[1]

Knowledge Transfer

The exact score transfers from continuous normal location to discrete Bernoulli proportion because both supply the same five roles: scalar regular likelihood, unbiased estimator, Fisher-information lower bound, actual variance and ratio. The likelihood algebra changes, but the recognition and comparison procedure is literal. Transferring to an irregular uniform-support model fails at the regularity role. Transferring to tests or experimental designs changes the measured object and performance criterion; the common word “efficiency” is insufficient.[1]

This narrower statistical calculation is not itself a prime. A general performance-versus-bound ratio across unrelated domains could be a future-prime question, but would need independent evidence and a stable boundary. The live prime Efficiency uses feasible-set dominance rather than merely division by a theorem lower bound, so this draft does not assume they are identical.

Examples

Normal location, all observations versus one. Let \(X_1,\ldots,X_n\) be IID \(N(\mu,\sigma^2)\) with known \(\sigma^2>0\), and take \(n\ge2\). The information for \(\mu\) is \(I_n=n/\sigma^2\), making the regular lower bound \(\sigma^2/n\). The sample mean is unbiased with variance \(\sigma^2/n\), so \(e_\mu(\bar X)=1\) for all \(\mu\). The deliberately wasteful estimator \(X_1\), though still unbiased when \(n\) observations are available, has variance \(\sigma^2\) and score \(1/n\). The comparison holds the experiment fixed; using one observation from a one-observation experiment would have a different information bound.[1]

Mapped back: The model/target are known-variance normal location and \(\mu\); the estimator/sample are \(\bar X\) or \(X_1\) computed from the same \(n\)-observation experiment; the information bound is \(\sigma^2/n\); actual variances are \(\sigma^2/n\) and \(\sigma^2\); and the bound-to-variance ratios are $1$ and \(1/n\).

Bernoulli proportion, all trials versus one. Let \(X_i\) be IID Bernoulli\((p)\) with $0<p<1$. For \(n\ge2\), \(I_n(p)=n/[p(1-p)]\) and the regular lower bound is \(p(1-p)/n\). The sample proportion is unbiased, has exactly that variance and scores one. Using only \(X_1\) from the same set of available trials is unbiased but has variance \(p(1-p)\), so scores \(1/n\). The case is discrete and has a different likelihood and parameter meaning from normal location, yet preserves the same structural comparison.[1]

Mapped back: The model/target are Bernoulli trials and success probability \(p\); the estimator/sample are \(\hat p=n^{-1}\sum_iX_i\) or \(X_1\) from the same \(n\) trials; the information bound is \(p(1-p)/n\); actual variances are \(p(1-p)/n\) and \(p(1-p)\); and ratios are $1$ and \(1/n\).

Structural Tensions

Sharp regular bound versus wider model coverage. Parameter-independent support and differentiability let the simple information inequality supply a powerful benchmark, but they exclude important models such as Uniform\([0,\theta]\). Pretending the benchmark applies there falsely ranks estimators; allowing the wider model class requires a different proof and may lose the same compact ratio. Diagnostic: Does the support or interchange step depend on \(\theta\), and has the claimed information bound actually been proved for this model?[1]

Unbiased precision versus total squared-error performance. Restricting to unbiased estimators makes variance the full MSE and permits this ratio. Allowing biased estimators can reduce variance but incurs squared bias, so their ranking under variance alone may disagree with ranking under MSE. Preserving the narrow score sacrifices coverage of potentially useful biased rules; broadening the decision class sacrifices the score's simple theorem interpretation. Diagnostic: Is the estimator unbiased in the stated model, and is the decision criterion variance among unbiased rules or total MSE across biased rules?[2]

Structural–Framed Character

This entry is structural inside a narrowly statistical frame. Evaluative weight is limited in the formula but re-enters through choosing variance and unbiasedness as the preferred performance criterion; another decision problem may value MSE or robustness. Human-practice dependence is low for the mathematical ratio once model and estimator are declared, but choosing those models and criteria is an analytical practice. Institutional origin is low: the construction is a theorem-based statistic, not a policy or organizational rule. Vocabulary travels literally between normal and Bernoulli estimation, while “efficiency” alone travels far beyond this formula and must not silently carry it. Import versus recognition is narrow: one may import the ratio to another regular scalar estimation model only after rechecking its likelihood, information and unbiasedness; otherwise the similarity is merely a benchmark analogy. Its character: a reusable formal precision measure within estimation theory, not a prime, because the Fisher-information and unbiased-sampling commitments are constitutive.[1][2]

Structural Core vs. Domain Accent

The broad skeleton is comparing observed performance with a justified benchmark. Here the benchmark is specifically inverse Fisher information and the observed performance is specifically sampling variance of an unbiased estimator. Live Estimator and Fisher information are strict prerequisites. Live Efficiency is a related frontier-based idea, not an asserted parent: \(1/I_n\) may be unattainable and therefore need not be a feasible alternative in the prime's no-dominated-slack sense. If a domain-neutral bound-to-performance ratio deserves a prime of its own, that is an explicit future-prime question, not an identity assumed by this draft.[1]

This entry presupposes Estimator and presupposes Fisher information. This score rates the variance of a specified statistical estimator. The score's regular-model benchmark is inverse Fisher information.

Relationships to Other Abstractions

Local relationship map for Cramér–Rao Estimator EfficiencyParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Cramér–RaoEstimator EfficiencyDOMAINDomain-specific abstraction: Estimator — presupposesEstimatorDOMAINDomain-specific abstraction: Fisher information — presupposesFisherinformationDOMAIN

Current abstraction Cramér–Rao Estimator Efficiency Domain-specific

Parents (2) — more general patterns this builds on

  • Cramér–Rao Estimator Efficiency presupposes Estimator Domain-specific

    This score rates the variance of a specified statistical estimator.

  • Cramér–Rao Estimator Efficiency presupposes Fisher information Domain-specific

    The score's regular-model benchmark is inverse Fisher information.

Hierarchy paths (2) — routes to 2 parentless roots

  • Cramér–Rao Estimator Efficiency → Estimator

Neighborhood in Abstraction Space

Cramér–Rao Estimator Efficiency sits in a sparse region of the domain-specific corpus (73rd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Statistical Learning & Model Failure Modes (41 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

The requested Wikipedia titles “Efficient estimator” and “Relative efficiency” both redirect to one broad page, but this narrower score does not settle either full identity. “Efficient estimator” is a property/object sense that may require equality across all parameter values; “relative efficiency” compares two procedures with a declared orientation and can exceed one. The \(2/\pi\) normal mean/median number is an asymptotic relative comparison, not a generic finite-sample score. Wilcoxon-versus-\(t\) compares tests and belongs to a different performance criterion. Experimental-design efficiency likewise needs its own design objective. The redirect and shared adjective are provenance, not alias or coverage proof.[3][2]

References

[1] Anna Mikusheva, 14.381 Statistical Method in Economics, Lecture 6: Efficient Estimators, Rao-Cramer Bound, MIT OpenCourseWare original instructor notes (2018), PDF pp.3–6: Fisher information, regularity, Theorem 3, Bernoulli equality and Uniform support counterexample. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r

[2] STATS 200: Introduction to Statistical Inference, Lecture 29, Stanford original course notes, PDF slides 16–20: bias–variance–MSE distinction and asymptotic efficiency. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k

[3] Charles J. Geyer, Statistics 5102 course slides, University of Minnesota original instructor material, PDF slides 48–56 on asymptotic relative efficiency and 244–245 on test comparisons. registry ↩a ↩b ↩c