Phi Coefficient¶
The signed Pearson correlation of two varying binary variables, computed from the normalized cross-product difference of their 2×2 table.
Core Idea¶
The phi coefficient (\(\phi\)) measures the signed association of two paired, varying binary variables. Arrange their observations in a \(2\times2\) table: \(a=n_{11}\) counts pairs where both are 1, \(b=n_{10}\) where the first is 1 and the second 0, \(c=n_{01}\) where the first is 0 and the second 1, and \(d=n_{00}\) where both are 0. With all four row and column margins positive, the statistic is
It is precisely the ordinary Pearson correlation of the two 0/1 columns. The numerator contrasts concordant and discordant cross-products; the denominator normalizes by variation in both binary variables. The sign records which cells dominate under the declared 0/1 coding. Switching just one variable's 0/1 labels reverses the sign, while switching both preserves it. The defined value lies in \([-1,1]\); if any marginal factor is zero, one variable is constant and mathematical phi is undefined.[1][2]
Phi describes the observed pairings. It is not a \(p\)-value, a causal effect or a guarantee of population dependence. In binary classification, the identical formula appears as the binary Matthews correlation coefficient (MCC) with \(a=TP\), \(b=FN\), \(c=FP\), \(d=TN\) when rows encode truth and columns encode prediction. A multiclass MCC extends beyond this \(2\times2\) identity.[2]
Structural Signature¶
Sig role-phrases: paired binary labels → four mutually exclusive joint cells → positive row and column margins → signed \(ad-bc\) contrast normalized by all margins → sample binary correlation under explicit coding.
- Paired carrier. The two labels must refer to the same observations. Comparing two separate prevalence totals supplies no joint cells.[1]
- Fourfold partition. Every pair occupies exactly one of $11\(, \$10\), $01\(, \$00\). The four cells produce both variables' marginal totals.[3][2]
- Signed contrast. \(ad-bc\) preserves association direction. Absolute value or squaring loses that direction and is not the same signed coefficient.[2]
- Four-factor normalization. \((a+b)(c+d)\) encodes the first variable's two margins; \((a+c)(b+d)\) encodes the second's. None may vanish.[2]
- Interpretation layer. Inference requires a separate sampling or null model; the fourfold statistic alone does not make a causal or population claim.[3]
What It Is Not¶
Phi is not a new alternative to Pearson's \(r\) for the same binary data. Compute Pearson correlation on the paired 0/1 columns, and the result is algebraically identical when both columns vary. What earns a separate specialist entry is its fourfold-table expression, coding behavior and diagnostic marginal constraints. It is likewise not an odds ratio: an odds ratio uses \(ad/bc\) when defined and is not restricted to \([-1,1]\).[1][2]
It is not unsigned association strength. Squared phi and related positive measures erase whether two 1-labels co-occur more or less than the margins imply. SciPy identifies Cramér's \(V\) and Tschuprow's \(T\) as related extensions, not interchangeable names for signed phi. The historically suggested “Doolittle Skill Score” alias remains unresolved: historical scores may be squared or use another convention, so it is deliberately absent from this entry's aliases.[4]
It is not every statistic computed from a confusion matrix. Youden's \(J\) combines sensitivity and specificity; Fowlkes–Mallows uses another precision/recall combination. Binary MCC is an exact formula match, but the multiclass MCC generalization is not this \(2\times2\) coefficient. Nor is phi automatically a superior classifier score because one class is rare: attainable values and decisions depend on margins, use and loss structure.[2]
Scope of Application¶
Any two binary attributes measured on the same units can form the table: an item may belong to one of two production groups and have one of two observed outcomes, or a classifier's prediction can be compared with its binary ground truth. H. P. Edmundson's original event-correlation treatment derives a correlation of events from indicator functions and relates it to classical Pearson correlation, providing the mathematical bridge between such settings.[1]
NIST's engineering-statistics handbook displays a hypothetical two-process by two-outcome table with observed cells $2,5;3,2$. It uses that table to teach Fisher's exact test. Applying phi separately to the displayed counts gives \(-11/35\approx-0.314\) under the row/column coding used here; NIST does not itself claim this phi result. The same data can support a descriptive coefficient and a separately specified test, but they answer different questions.[3]
Scikit-learn's original metric guide gives a four-case truth/prediction illustration: true labels \((+,+,+,-)\) and predicted labels \((+,-,+,+)\). The corresponding cells are \(TP=2\), \(FN=1\), \(FP=1\), \(TN=0\), yielding \(-1/3\) (reported approximately \(-0.33\)). The zero true-negative cell does not make phi undefined because all row and column margins remain positive.[2]
Clarity¶
The cell orientation must be declared. With \(a=11\), \(b=10\), \(c=01\), \(d=00\), the sign follows \(ad-bc\). Exchanging only the codes of the second variable exchanges the corresponding columns and reverses the sign. In a prediction table, calling a cell “true positive” depends on which label was designated positive; the signed interpretation is not invariant to an unreported coding switch.[2]
The four margins are as important as the four cells. If every observation is in one row or one column, the denominator is zero and Pearson correlation is undefined. A software library may choose a numerical fallback for convenience, but that implementation convention should not be mistaken for the mathematical coefficient. By contrast, \(d=0\) alone in scikit-learn's four-case example leaves a valid nonzero denominator.[2]
Zero phi means zero observed binary correlation in the table. With positive margins in a complete \(2\times2\) distribution, the zero determinant also expresses factorization of those empirical proportions. It does not establish that the data-generating population is independent or that no causal relation is possible; such conclusions require evidence and assumptions beyond the table.[1][3]
Manages Complexity¶
Phi compresses four joint-cell counts into a single signed, normalized comparison. This makes the direction of observed association legible and permits a quick algebraic check against Pearson correlation computed from raw binary rows. It also exposes errors: forgetting one marginal factor or swapping a cell changes the normalization or sign.[1][2]
The compression hides margins. For a fixed sample size \(n\) and row/column positive totals \(r_1=a+b\) and \(c_1=a+c\), the possible \(a\) values satisfy \(\max(0,r_1+c_1-n)\le a\le\min(r_1,c_1)\). Since \(ad-bc=na-r_1c_1\), phi changes monotonically with \(a\) when those margins are fixed. The general mathematical range is \([-1,1]\), but fixed margins can block an endpoint. A score without its margins conceals that feasibility constraint.[3][2]
Abstract Reasoning¶
First verify that the observations are paired: each unit has a value for both binary variables. Declare the 1/0 coding, tabulate \(a,b,c,d\), compute all row and column margins and reject a zero denominator. Calculate the signed cross-product difference, divide by the four-factor square root and compare with Pearson \(r\) on the 0/1 columns as a consistency check.[1][2]
Then choose the interpretation appropriate to the task. A descriptive comparison can stop at phi and the table. A classifier report may call the same number binary MCC; a significance test requires a separate null and sampling plan; a causal argument requires far more. NIST's Fisher example demonstrates why a fourfold table and a test statistic must not be collapsed into a single unnamed verdict.[3]
Knowledge Transfer¶
The two worked settings preserve four roles: paired binary labels, four joint cells, signed cross-product difference and normalization by both variables' margins. “Process group versus outcome” and “prediction versus truth” do not have the same substantive interpretation, but the phi computation is literal in both. Under the second labeling, the exact binary MCC expression is a recognized application of the same statistic.[3][2]
Do not transfer the task-specific verdict. A negative phi in the NIST hypothetical says that the chosen process-row code and plus-outcome code are inversely associated in those displayed counts; the scikit-learn example concerns errors in four predictions. Neither value alone establishes causality, future performance or statistical significance, and a multiclass classifier requires a different extension.[3][2]
Examples¶
Process group and binary outcome. NIST's displayed hypothetical table has row 1 counts \((2,5)\) and row 2 counts \((3,2)\). Let first variable 1 mean row 1 and second variable 1 mean the first outcome column. Then \((a,b,c,d)=(2,5,3,2)\) and \(\phi=(4-15)/\sqrt{7\cdot5\cdot5\cdot7}=-11/35\). This is a derived phi value from NIST's counts, not a phi calculation NIST reports.[3][2]
Mapped back: paired observations = process row/outcome column for each item; fourfold counts = $2,5,3,2$; signed contrast = \(-11\); positive margins = $7,5,5,7$.
Binary classifier. In scikit-learn's original example, two true positives, one false negative, one false positive and zero true negatives give \(\phi=MCC=(2\cdot0-1\cdot1)/\sqrt{3\cdot1\cdot3\cdot1}=-1/3\). The documented output rounds this to \(-0.33\).[2]
Mapped back: paired observations = truth/prediction per case; fourfold counts = \(TP,FN,FP,TN\); signed contrast = \(-1\); positive margins = $3,1,3,1$ despite one empty cell.
Structural Tensions¶
Standardized score versus marginal feasibility. Normalization places defined phi on a common signed Pearson scale, but unequal fixed row and column totals constrain which fourfold tables are possible. In NIST's \(7/5\) row and \(5/7\) column example, the fixed margins permit \(a=0\) through $5$: the minimum reaches \(-1\), while the maximum is only \(5/7\). It is therefore wrong to say that both extremes are always narrowed whenever the margins differ.[3][2]
Diagnostic: Before ranking two observed scores, do their respective margins allow the same attainable positive and negative extremes?
Descriptive compression versus inferential warrant. A compact signed number states how the observed binary labels covary. A null test, interval or causal explanation requires an additional design. NIST's Fisher procedure supplies one separate conditional test for its table, not an intrinsic \(p\)-value attached to phi.[3]
Diagnostic: Is the claim about this table's correlation, or has a particular sampling/null/causal assumption been supplied for the stronger claim?
Structural–Framed Character¶
Its character: phi lies near the structural end of the spectrum: it is an exact mathematical specialization of Pearson correlation. Its applied frames—quality comparison, classification, event association—determine label meaning and evaluation goals, not the formula itself.
- Vocabulary travels: the fourfold calculation survives the move from industrial comparisons to prediction assessment.[3][2]
- Evaluative weight: the signed value is descriptive; whether a sign is “good” depends on coding and task objective.
- Institutional origin: historical statistical and contemporary software communities name the coefficient, but a name does not alter its algebra.[1][2]
- Human-practice dependence: researchers select variables, coding, sample and whether they report inference.
- Import versus recognition: once two binary columns and nonzero margins exist, the correlation can be recognized mathematically; calling any unrelated four-cell score “phi” would improperly import this normalization.
Structural Core vs. Domain Accent¶
The portable core is normalized covariance of two nonconstant variables, already represented by live Pearson correlation coefficient and ultimately the prime Correlation. Phi adds the binary carrier: four joint counts, all four marginal factors, sign under label swaps and fixed-margin feasibility. It remains a domain-specific statistics node because it does not establish a separate cross-substrate prime; it is a precise species of an existing statistic.[1][2]
The proposed strict Pearson parent is warranted algebraically, unlike a merely topical association edge. Youden's \(J\) and Fowlkes–Mallows are neighboring confusion-matrix summaries with different normalizations or targets, not genuses of signed phi. No Doolittle alias is applied while its historical formula and naming remain unresolved.
Instantiates / Related Primes¶
This entry is a kind of Pearson correlation coefficient.
Correlation is an upstream general relation already represented above Pearson. The binary MCC label describes an equal formula in the classifier setting, while multiclass MCC and Cramér's \(V\) remain separate extensions.[1][2][4]
Relationships to Other Abstractions¶
Current abstraction Phi Coefficient Domain-specific
Parents (1) — more general patterns this builds on
-
Phi Coefficient is a kind of Pearson correlation coefficient Domain-specific
Pearson correlation specialized to two binary variables.Encoding paired binary observations as 0/1 makes their Pearson covariance divided by both standard deviations algebraically equal the signed 2×2 phi formula. The live Pearson coefficient supplies the genus; a binary carrier with fourfold-table margins supplies the species differentia.
Hierarchy path (1) — routes to 1 parentless root
- Phi Coefficient → Pearson correlation coefficient → Correlation
Neighborhood in Abstraction Space¶
Phi Coefficient sits in a sparse region of the domain-specific corpus (76th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Clinical Trial & Research Methodology (20 abstractions)
Nearest neighbors
- Quadrant Count Ratio — 0.84
- Pair Distribution Function — 0.83
- Join Count Statistic — 0.83
- Log-Linear Analysis — 0.83
- Covariance Matrix — 0.82
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
Odds ratio is unbounded when defined and does not normalize by four margins; phi-squared discards the sign; Pearson contingency coefficient is a different nonnegative statistic despite the Pearson name; Cramér's \(V\) is a related unsigned association measure; Fisher's exact test evaluates a null distribution rather than merely returning this correlation; and multiclass MCC is not the binary fourfold coefficient. “Doolittle Skill Score” is held for separate lexical/historical adjudication rather than presumed synonymous.[3][2][4]
References¶
[1] H. P. Edmundson, “A Correlation Coefficient for Attributes or Events,” in Statistical Association Methods for Mechanized Documentation, original National Bureau of Standards symposium proceedings (1964), printed p.41 onward, Abstract and §3 formula (3.1). The full PDF exceeded the web viewer's size limit; only official indexed original-source passages were used, and no disputed historical phi-squared wording is generalized. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j
[2] Scikit-learn maintainers, “Metrics and scoring,” §3.4.4.13, Matthews correlation coefficient, original documentation, binary formula, multiclass extension and four-label worked output. The guide quotes Wikipedia for the historical name; it is used here for the maintainer-described formula and example. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w
[3] NIST/SEMATECH, Engineering Statistics Handbook, §7.3.3, original handbook's two-process by binary-outcome table and Fisher-test illustration, especially observed cells \((2,5;3,2)\). Its authors do not compute phi there; values shown here are explicit algebraic derivations. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m
[4] SciPy maintainers, scipy.stats.contingency.association, original documentation, Notes on Cramér's \(V\) and Tschuprow's \(T\) as related extensions. registry ↩a ↩b ↩c