Skip to content

Sargan–Hansen Test

An overidentification test for instrumental-variable or GMM models that asks whether surplus instruments are jointly orthogonal to fitted residual moments, conditional on at least one maintained valid identifying set.

Version
v1 · 2026-08-30 · History
Domain-specific #
2707
Origin domain
economics
Subdomain
econometrics and generalized method of moments
Aliases
Test of overidentifying restrictions, Hansen J test, Sargan test

Core Idea

The Sargan–Hansen family tests the joint moment restrictions left over when an instrumental-variable or GMM model has more instruments than endogenous parameters to identify. After estimating the model, it measures how far sample moments between instruments and residuals depart from zero. Under the maintained model and regularity conditions, the optimized GMM criterion produces a chi-squared statistic with degrees of freedom equal to the number of overidentifying restrictions.[1]

The classical Sargan form is tied to homoskedastic linear IV assumptions; Hansen's J statistic uses an appropriate GMM weighting/covariance estimate and can be robust to heteroskedasticity under its conditions. Rejection says the restrictions are jointly incompatible with the data and maintained specification. Nonrejection does not prove every instrument valid, and the test cannot operate in an exactly identified model. Weak instruments, many instruments, dependence, misspecified covariance, or failure of all candidate instruments can distort interpretation.[2]

Structural Signature

  • The structural model. Parameters and endogenous regressors are declared.
  • The instrument set. Excluded or included instruments supply candidate moments.
  • The moment restrictions. Validity implies zero population correlation with structural disturbance.
  • The overidentification count. There are more independent moments than estimated parameters.
  • The fitted residuals. Estimation supplies sample violations.
  • The weighting matrix. Assumptions determine how moments are scaled and combined.
  • The quadratic statistic. The minimized criterion summarizes joint discrepancy.
  • The reference distribution. Asymptotic chi-square calibration uses the surplus restriction count.
  • The diagnostic conclusion. Rejection challenges the joint instrument/model package, not one instrument uniquely.

What It Is Not

  • Not an instrument-relevance test. First-stage strength is a separate question.
  • Not proof of exogeneity after nonrejection. Low power and common invalidity remain possible.
  • Not available under exact identification. No surplus restriction remains to test.
  • Not a selector of the guilty instrument. The null is joint.
  • Not automatically heteroskedasticity-robust. Statistic and covariance construction must match assumptions.
  • Not a substitute for causal design. Institutional knowledge and exclusion arguments remain necessary.

Scope of Application

The test is literal in IV, two-stage least squares, panel IV, and GMM model diagnostics.

  • Linear IV. Checking surplus exclusion restrictions under stated error assumptions.
  • GMM. Testing the full vector of overidentifying moments.
  • Panel estimators. Auditing growing lag-instrument sets while guarding against proliferation.
  • Causal studies. Reporting a limited falsification check alongside design arguments.
  • Specification comparison. Seeing whether added instruments make moments jointly inconsistent.
  • Simulation. Studying size and power under weak or invalid instruments.
  • Replication. Verifying degrees of freedom, weighting, and covariance choices.

Clarity

State model, endogenous variables, every instrument and exclusion rationale, moment vector, estimator, weighting matrix, covariance assumptions, clustering/dependence adjustment, statistic, degrees of freedom, sample size, and p-value. Report strength diagnostics and instrument count separately. Interpret rejection as joint misspecification and nonrejection as failure to reject, never validation.

State the estimator, moment conditions, weighting matrix, number of parameters identified by instruments, and number of overidentifying restrictions. The test exists only when surplus moment restrictions remain; an exactly identified model has zero degrees of freedom and no overidentification test. The null is joint validity of the tested moments conditional on the maintained model and at least enough valid moments for identification. Rejection does not identify which instrument is invalid or whether misspecification instead lies in the structural equation, functional form, or error assumptions. The classical Sargan statistic and robust Hansen J statistic use different covariance conditions and should not be treated as interchangeable names in every estimator. Weak instruments, many instruments, clustering, heteroskedasticity, and finite samples can distort interpretation even when software reports a familiar chi-squared reference.

Manages Complexity

The test compresses many moment discrepancies into one calibrated statistic and uses surplus identifying information as an internal consistency check. That convenience sacrifices localization and depends on asymptotics. Instrument proliferation can yield deceptively weak tests, while a rejected statistic may reflect structural, dynamic, covariance, or exclusion misspecification.

Instrumental-variable models can fit coefficients while leaving unresolved whether the instruments are orthogonal to the structural disturbance. Overidentification turns surplus instruments into testable restrictions: the fitted residual moments cannot all be freely set to zero once parameters have been estimated. The optimized GMM objective aggregates their standardized discrepancy into one statistic. This compresses a vector of moment failures, but the compression loses localization. A small value can reflect genuinely valid moments or low power, and a large value can reflect one bad instrument, correlated misspecification, or a poor covariance estimate. Difference tests and instrument-subset analysis can refine diagnosis only under their own maintained conditions. The abstraction manages complexity by testing model-implied residual orthogonality, not by certifying instruments individually.

Abstract Reasoning

  1. Specify parameters and population moment restrictions.
  2. Check that moments outnumber estimated parameters.
  3. Estimate the model with the declared IV/GMM procedure.
  4. Compute residual sample moments.
  5. Choose a valid weighting and covariance estimate.
  6. Form the minimized quadratic statistic and degrees of freedom.
  7. Compare with the reference distribution.
  8. Diagnose the joint package and perform separate strength/sensitivity analyses.

Knowledge Transfer

The test exemplifies hypothesis testing: a maintained restriction generates a reference distribution for a discrepancy statistic. Its special leverage comes from redundant moment conditions; the null/alternative parent travels, while instruments and GMM geometry keep this domain-specific.

Null-versus-Alternative Hypothesis Testing is the strict parent because the statistic calibrates observed moment discrepancy against a null distribution and yields a reject-or-not decision at a chosen level. The transferable pattern is reserve redundant constraints → fit with necessary constraints → test whether the surplus constraints remain compatible. It appears in specification testing beyond econometrics, but the Sargan–Hansen residual includes instruments, GMM moments, optimized weighting, and degrees of freedom from overidentification. It is not a weak-instrument diagnostic and cannot replace substantive instrument justification.

Examples

Canonical

A linear model has one endogenous regressor and three excluded instruments. After estimation, two overidentifying restrictions remain. The joint residual–instrument moment discrepancy is compared with a chi-square distribution with two degrees of freedom under the applicable Sargan or Hansen assumptions.[1]

Mapped back: surplus valid moments → residual orthogonality predictions → weighted discrepancy → joint specification test.

Applied / In Practice

A panel GMM study reports a very high J-test p-value with almost as many instruments as units. Rather than calling this confirmation, the analyst collapses/reduces instruments, checks serial correlation and strength, and reports sensitivity because instrument proliferation can weaken the diagnostic.

A model uses more excluded instruments than are needed to identify one endogenous regressor. After estimation, the analyst computes the jointly weighted correlation between instruments and fitted residuals. A rejection prompts investigation of instrument subsets, structural specification, and covariance assumptions, not immediate deletion of whichever instrument has the largest raw correlation. A failure to reject is reported with power and weak-instrument diagnostics rather than as proof of exogeneity. If the model is re-estimated with exactly enough instruments, the overidentification statistic disappears because no surplus restriction remains. The example makes the conditional null and degrees-of-freedom logic visible.

Mapped back: nominal nonrejection + weak diagnostic regime → instrument-set sensitivity → qualified inference.

Structural Tensions

  • Redundant identification vs. additional assumptions. Extra instruments create a test by adding claims that can fail. Diagnostic: What exclusion does each instrument require?
  • Joint power vs. localization. One statistic detects inconsistency but not its source. Diagnostic: Which restricted subsets and specifications reproduce rejection?
  • Robustness vs. finite-sample distortion. Asymptotic corrections do not solve weak/many-instrument problems. Diagnostic: Is the sample regime adequate?
  • Nonrejection vs. validation. Low power can mimic fit. Diagnostic: What alternatives could the test detect?
  • Autonomous test vs. generic hypothesis testing. Many tests use chi-square statistics; surplus IV/GMM moments define this identity. Diagnostic: Are overidentifying restrictions the tested object?

Structural–Framed Character

The test is structural under a statistical model. Sample moments and computation are objective, while instrument validity, covariance, asymptotics, and significance thresholds frame inference. It is evaluatively neutral but causally consequential. Hypothesis Testing supplies the logic; GMM supplies the geometry.

Estimated moment model, surplus orthogonality restrictions, residual-moment vector, covariance weighting, reference distribution, and conditional joint-validity interpretation are structural. Instrument labels, dataset, software command, covariance estimator, sample size, and significance threshold are framed. Robustness choices can change which named version is appropriate. The result is conditional on the maintained identifying set and model specification; those are not proven by failure to reject. This framing keeps a convenient scalar p-value from becoming an unconditional validity certificate.

Structural Core vs. Domain Accent

The skeleton is maintained restrictions → discrepancy statistic → calibrated rejection decision. The accent is endogenous regressors, instruments, residual orthogonality, surplus moments, GMM weighting, and chi-square degrees. Remove those and one has hypothesis testing generally.

The portable core is use constraints beyond those required for fitting as a specification check. The econometric accent is instrumental-variable or GMM estimation, residual orthogonality moments, an optimized quadratic criterion, and degrees of freedom equal to surplus restrictions. Remove those roles and the method becomes hypothesis testing generally. Keep the moments but lack overidentification, and there is no Sargan–Hansen test to compute. The autonomous residual also includes its conditional interpretation: the statistic tests the joint compatibility of surplus moments under maintained assumptions rather than discovering a uniquely invalid instrument. Instrument relevance and finite-sample calibration remain separate diagnostic questions.

Hypothesis Testing (Null vs. Alternative) is the strict parent because the statistic evaluates a null of joint moment validity against misspecification. The parent applies far beyond instrumental variables.

The prospective workspace queue contains one strict upward edge to prime:hypothesis_testing_null_vs_alternative. No live DAG mutation is authorized.

Relationships to Other Abstractions

Local relationship map for Sargan–Hansen TestParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Sargan–Hansen TestDOMAINPrime abstraction: Hypothesis Testing (Null vs. Alternative) — is a kind ofHypothesis Test…PRIME

Current abstraction Sargan–Hansen Test Domain-specific

Parents (1) — more general patterns this builds on

  • Sargan–Hansen Test is a kind of Hypothesis Testing (Null vs. Alternative) Prime

    Hypothesis Testing (Null vs.

Neighborhood in Abstraction Space

Sargan–Hansen Test sits in a sparse region of the domain-specific corpus (87th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Causal Identification & Endogeneity (9 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-09-08

Not to Be Confused With

  • Weak-instrument test. Evaluates relevance/identification strength.
  • Hausman test. Compares estimators under different consistency claims.
  • Durbin–Wu–Hausman endogeneity test. Tests regressor exogeneity.
  • Difference-in-Hansen test. Tests a subset by differencing nested J statistics.
  • Sargan test. The homoskedastic linear-IV member of the family.
  • Hansen J test. The GMM formulation under its covariance assumptions.

References

[1] Lars Peter Hansen, ‘Large Sample Properties of Generalized Method of Moments Estimators,’ Econometrica 50, no. 4 (1982): 1029–1054, https://doi.org/10.2307/1912775. registry ↩a ↩b

[2] J. D. Sargan, ‘The Estimation of Economic Relationships Using Instrumental Variables,’ Econometrica 26, no. 3 (1958): 393–415, https://doi.org/10.2307/1907619. registry