Hat matrix¶
The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation.
Core Idea¶
Hat matrix is treated here as the recurring mathematics and formal science identity summarized by this source-grounded definition: The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation.
In statistics, the projection matrix (\mathbf{P}) , sometimes also called the influence matrix or hat matrix (\mathbf{H}) , maps the vector of response values (dependent variable values) to the vector of fitted values (or predicted values). It describes the influence each response value has on each fitted value. The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation.
(Note that \left( \mathbf{X}^\textsf{T} \mathbf{X} \right)^{-1} \mathbf{X}^\textsf{T} is the pseudoinverse of X.) Some facts of the projection matrix in this setting are summarized as follows. If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,. As \mathbf{\hat{y}} is usually pronounced "y-hat", the projection matrix \mathbf{P} is also named hat matrix as it "puts a hat on \mathbf{y} ".
For Hat matrix, the abstraction is narrower than the article's general subject matter: a positive case must preserve The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation. Retaining only the name, a familiar example, or a downstream effect is insufficient. The specialist roles and tests remain anchored in mathematics and formal science, which is why this identity is domain-specific rather than prime.
Structural Signature¶
Sig role-phrases:
- Defining carrier — If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,.
- Constitutive relation — The covariance matrix of the residuals \mathbf{r} , by error propagation, equals.
- Operating condition — where \mathbf{\Sigma} is the covariance matrix of the error vector (and by extension, the response vector as well).
- Recognition evidence — Suppose the design matrix \mathbf{X} can be decomposed by columns as \mathbf{X} = \begin{bmatrix} \mathbf{A} & \mathbf{B} \end{bmatrix} .
- Admissible variation — In the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression.
- Characteristic consequence — The hat matrix was introduced by John Wilder in 1972.
- Failure boundary — An article by Hoaglin, D.C. and Welsch, R.E.
What It Is Not¶
- Not the whole field of mathematics and formal science. The node requires the specific identity stated by The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation.
- Not an over-broad reading. However, this is not always the case; in locally weighted scatterplot smoothing (LOESS), for example, the hat matrix is in general neither symmetric nor idempotent.
- Not an over-broad reading. The above may be generalized to the cases where the weights are not identical and/or the errors are correlated.
- Not an over-broad reading. If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,.
- Not automatically Matrix variate Dirichlet distribution. Retrieval proximity does not establish equivalence; the two identities must be compared by carrier, operation, and failure boundary.
Scope of Application¶
Hat matrix applies literally inside mathematics and formal science wherever the source-defined carrier and relation can be established. Its documented habitats include:
- Blockwise formula. In the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression.
- Properties. For other models such as LOESS that are still linear in the observations \mathbf{y} , the projection matrix can be used to define the effective degrees of freedom of the model.
- Properties. Practical applications of the projection matrix in regression analysis include leverage and Cook's distance, which are concerned with identifying influential observations, i.e. observations which have a large effect on the results of a regression.
- History. (1978) gives the properties of the matrix and also many examples of its application.
- Blockwise formula. There are a number of applications of such a decomposition.
- Definition. If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,.
Outside mathematics and formal science, the name should be retained only when these same operational conditions survive; otherwise the comparison belongs to the broader parent Pattern or should be marked as analogy.
Clarity¶
A clear use of Hat matrix names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation. The strongest recognition evidence in the frozen account is: Suppose the design matrix \mathbf{X} can be decomposed by columns as \mathbf{X} = \begin{bmatrix} \mathbf{A} & \mathbf{B} \end{bmatrix} . A report should distinguish that evidence from a proxy, consequence, or common implementation. It should also state the qualification However, this is not always the case; in locally weighted scatterplot smoothing (LOESS), for example, the hat matrix is in general neither symmetric nor idempotent. so that a reader can reproduce the classification rather than infer it from topical resemblance.
Manages Complexity¶
Hat matrix compresses multiple mathematics and formal science details into a stable diagnostic relation. The source shows both the central mechanism—the covariance matrix of the residuals \mathbf{r} , by error propagation, equals.—and the practical consequence—the hat matrix was introduced by John Wilder in 1972. This compression makes cases comparable while leaving parameters, conventions, exceptions, and evidential quality explicit. It is lossy by design: local history and implementation details may be omitted only when they do not alter the defining relation.
Abstract Reasoning¶
- Type the carrier. Identify the mathematics and formal science entities to which the claim applies.
- State the relation. Use the source-grounded identity: The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation.
- Check operation and conditions. where \mathbf{\Sigma} is the covariance matrix of the error vector (and by extension, the response vector as well).
- Demand recognition evidence. Suppose the design matrix \mathbf{X} can be decomposed by columns as \mathbf{X} = \begin{bmatrix} \mathbf{A} & \mathbf{B} \end{bmatrix} .
- Test variation. Change an implementation or setting while preserving in the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression.
- Run the collapse test. Remove the defining operation; if the label still seems equally apt, only a topic or correlate was retained.
- Reduce cautiously. When the specialist conditions cannot be carried, route the residual comparison to Pattern.
Knowledge Transfer¶
Within the home domain. Knowledge about Hat matrix transfers literally when a new case preserves the same carrier type, relation, and recognition test. In the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression. For other models such as LOESS that are still linear in the observations \mathbf{y} , the projection matrix can be used to define the effective degrees of freedom of the model.
Beyond the home domain. No canonical parent is asserted for Hat matrix. An outside case receives the specialist name only when the same typed roles and rejection conditions can be filled literally; otherwise the comparison remains an analogy pending later graph densification.
Examples¶
Canonical¶
However, this is not always the case; in locally weighted scatterplot smoothing (LOESS), for example, the hat matrix is in general neither symmetric nor idempotent. This case is canonical because it supplies a concrete carrier and lets the defining relation be checked rather than merely named.
Mapped back: carrier → the entities in the documented case; operation → The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation; recognition evidence → Suppose the design matrix \mathbf{X} can be decomposed by columns as \mathbf{X} = \begin{bmatrix} \mathbf{A} & \mathbf{B} \end{bmatrix}
Applied / In Practice¶
For the case of linear models with independent and identically distributed errors in which \mathbf{\Sigma} = \sigma^{2} \mathbf{I} , this reduces to. The applied case shows how the identity is used under a second setting or qualification while keeping the same operative relation.
Mapped back: changed setting → Application for residuals; invariant → The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation; boundary → the case exits the class when however, this is not always the case; in locally weighted scatterplot smoothing (LOESS), for example, the hat matrix is in general neither symmetric nor idempotent
Structural Tensions¶
T1 — Stable identity versus admissible variation. However, this is not always the case; in locally weighted scatterplot smoothing (LOESS), for example, the hat matrix is in general neither symmetric nor idempotent. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Which changes preserve the defining relation, and which replace it?
T2 — Recognition versus proxy. The above may be generalized to the cases where the weights are not identical and/or the errors are correlated. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Does the cited evidence establish the identity or only a correlated sign?
T3 — Definition versus implementation. If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Is the observed implementation constitutive, optional, or merely common?
T4 — Scope versus overextension. As \mathbf{\hat{y}} is usually pronounced "y-hat", the projection matrix \mathbf{P} is also named hat matrix as it "puts a hat on \mathbf{y} ". The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Can every claimed application fill the same typed roles without metaphor?
T5 — Transfer versus domain accent. If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Does the receiving case instantiate Hat matrix literally, co-instantiate Pattern, or only resemble it?
T6 — Autonomy versus reduction. The covariance matrix of the residuals \mathbf{r} , by error propagation, equals. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: What does Hat matrix distinguish that the broader parent Pattern leaves together?
Structural–Framed Character¶
Hat matrix is structural-leaning. Its structural side is the repeatable organization summarized by The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation. Its framed side is the mathematics and formal science vocabulary that fixes the carrier, evidence, exceptions, and admissible transformations.
Evaluative weight: the identity can be stated descriptively even when applications carry practical stakes. Human-practice dependence: the source-grounded carrier determines whether the relation exists independently or is constituted by a practice. Institutional origin: disciplinary conventions stabilize the name and test. Vocabulary portability: where \mathbf{\Sigma} is the covariance matrix of the error vector (and by extension, the response vector as well). Import versus recognition: literal transfer requires the same mechanism; shape alone is analogy.
Its portable skeleton is Pattern. Its character: a recurring specialist identity whose thin organization can be abstracted, while its operational meaning remains domain-bound.
Structural Core vs. Domain Accent¶
What is skeletal. The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation. The stable skeleton is the typed relation expressed in that definition and the entry's recognition and collapse tests. The source identifies these operative conditions: If the vector of response values is denoted by \mathbf{y} and the vector of fitted values by \mathbf{\hat{y}} ,. The covariance matrix of the residuals \mathbf{r} , by error propagation, equals. It further constrains recognition and variation through: where \mathbf{\Sigma} is the covariance matrix of the error vector (and by extension, the response vector as well). Suppose the design matrix \mathbf{X} can be decomposed by columns as \mathbf{X} = \begin{bmatrix} \mathbf{A} & \mathbf{B} \end{bmatrix} .
What is domain-bound. mathematics and formal science supplies the operative entities, technical vocabulary, warrants, and exceptions that make Hat matrix literal. Its documented scope includes the condition that In the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression. Another bounded application condition is that For other models such as LOESS that are still linear in the observations \mathbf{y} , the projection matrix can be used to define the effective degrees of freedom of the model. These are not decorative examples; they determine which carrier and evidence can fill the abstraction's roles.
Why no parent is asserted. Removing those specialist details does not currently yield one live catalog node that is a necessary genus for every instance. The entry is therefore approved as unparented rather than attached by topical resemblance. Its collapse evidence remains specific—In the classical application \mathbf{A} is a column of all ones, which allows one to analyze the effects of adding an intercept term to a regression.—and future graph densification may discover a defensible relation only if it preserves that boundary.
Instantiates / Related Primes¶
This entry is a kind of Matrix.
- Approved unparented node. No current live node supplies a defensible necessary genus or structural prerequisite for Hat matrix. The reviewed identity is: The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation. The accelerated suggestion was declined because topical or lexical similarity does not establish hierarchy; the node is admitted without a parent pending later graph densification.
- Related reasoning operations. Evidence, representation, comparison, classification, transformation, or evaluation may participate in particular cases, but participation does not make any one of them a necessary parent of every instance.
Relationships to Other Abstractions¶
Current abstraction Hat matrix Domain-specific
Parents (1) — more general patterns this builds on
-
Hat matrix is a kind of Matrix Domain-specific
A hat matrix is the projection matrix that maps observed responses to fitted values in linear regression.A hat matrix is the projection matrix that maps observed responses to fitted values in linear regression.
Hierarchy paths (5) — routes to 5 parentless roots
- Hat matrix → Matrix → Tensor → Transformation → Function (Mapping)
- Hat matrix → Matrix → Linearity
- Hat matrix → Matrix → Representation → Abstraction
- Hat matrix → Matrix → Tensor → Invariance
- Hat matrix → Matrix → Tensor → Vector Space → Set and Membership
Neighborhood in Abstraction Space¶
Hat matrix sits in a crowded region of the domain-specific corpus (32nd percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Unclustered & Miscellaneous (2551 abstractions)
Nearest neighbors
- S-procedure — 0.91
- Durbin–Wu–Hausman test — 0.89
- Single Vegetative Obstruction Model — 0.88
- Mathematical Modeling — 0.88
- Entropy estimation — 0.87
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Pattern. The parent omits the specialist differentia. Tell: Can the case establish The diagonal elements of the projection matrix are the leverages, which describe the influence each response value has on the fitted value for that same observation?
- Matrix variate Dirichlet distribution. A probability law on several positive-definite matrices whose sum remains below the identity, generalizing scalar Dirichlet and matrix beta distributions. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- Matrix population models. Stage- or age-structured population models that project abundance by multiplying a population-state vector by a matrix of survival, transition, growth, and reproduction rates. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- Moran's I. A weighted statistic measuring global spatial autocorrelation by comparing cross-products among neighboring observations with overall variance. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- A measurement, proxy, or consequence. Those may provide evidence without being the identity. Tell: Would Hat matrix remain present if the detector or downstream effect changed?
- A metaphorical analogue. A similar shape outside mathematics and formal science lacks the specialist mechanism. Tell: Do the native roles transfer literally, or only the parent Pattern?
References¶
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Projection_matrix (revision 1343363405).
- Preserved source candidate: https://books.google.com/books?id=ScssAwAAQBAJ&pg=PA160
- Preserved source candidate: http://old.ecmwf.int/newsevents/training/lecture_notes/pdf_files/ASSIM/ObservationInfluence.pdf
- Preserved source candidate: https://web.archive.org/web/20140903115021/http://old.ecmwf.int/newsevents/training/lecture_notes/pdf_files/ASSIM/ObservationInfluence.pdf
- Preserved source candidate: http://dspace.mit.edu/bitstream/1721.1/1920/1/SWP-0901-02752210.pdf
- Preserved source candidate: https://archive.org/details/datafittinginche0000gans
- Preserved source candidate: https://archive.org/details/advancedeconomet00amem/page/460
- Preserved source candidate: https://archive.org/details/advancedeconomet00amem
- Preserved source candidate: https://math.stackexchange.com/q/1582567
The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.