Skip to content

Counternull

A nonnull effect value or set that matches a designated null's p-value under a specified test of the observed data.

Version
v2 · 2026-10-03 · History
Domain-specific #
13101
Domain group
Formal Sciences
Origin domain
Experimental Design & Statistics
Subdomains
Hypothesis Testing, Effect Size Reporting → Experimental Design & Statistics
Aliases
Counternull Value, Counternull Set

Core Idea

A counternull is a nonnull effect or parameter value that, when tested against the observed data by the specified procedure, receives the same p-value as a designated null value. Robert Rosenthal and Donald Rubin introduced the counternull value as a companion to the usual null-hypothesis result: if that nonzero value were used as the tested null, the resulting p-value would match the one reported for the original null.[1] The construction asks a pointed question: besides zero, what magnitude of effect would this same test treat just as it treats zero? The answer can reveal why “not statistically significant” does not establish a zero effect, while a small p-value alone does not establish a scientifically important magnitude.

The equality is test-relative. It concerns tail probabilities computed from a specified statistic, model or randomization distribution, extremeness rule, and observed data. It does not by itself mean that the null and counternull have equal posterior probabilities, that both effects are equally credible under every evidential standard, or that either is practically important. Under a symmetric one-dimensional reference distribution, the observed estimate lies halfway between null and counternull. For a zero null this often gives the familiar shortcut “counternull = twice the observed estimate.” Rosenthal and Rubin explicitly limit that shortcut to suitable common cases; it is not the definition.[1][2]

The name covers a scalar value in the classical illustration and a possibly nonunique Η set under later randomization-based formulations. Bind and Rubin show that changing the test statistic can change that set and that the equal-p condition need not always produce a unique nonnull solution.[2] Thus one reports the tested null, test and statistic along with the counternull; the number alone is not an invariant property of a dataset.

Structural Signature

Sig role-phrases: designated null value → observed estimate or statistic → specified hypothesis-dependent test → nonnull candidate value or set → equal-p comparison.

  • Designated null value: A selected effect or parameter value, often zero, supplies the original test reference. A counternull is defined in relation to it, not in isolation.
  • Observed estimate or test statistic: The realized data fix the location whose extremeness is assessed under each candidate hypothesis. The counternull is specific to this observation and analysis.
  • Specified hypothesis-dependent test: A probability model or randomization scheme, statistic, direction of extremeness, and tail rule determine each p-value. Without them, “same support” is underspecified.
  • Nonnull candidate value or set: At least one effect or parameter value distinct from the designated null is considered. Depending on the procedure, there can be one, several, a continuum, or no matching nonnull values.[2]
  • Equal-p comparison: The defining filter retains candidates whose calculated p-value equals the null's calculated p-value under the declared comparison rule. Equal tail scores are the operational criterion; the intended interpretation is a separate issue.[1]

In notation, for data \(x\), chosen test construction \(T\), designated null \(\theta_0\), and hypothesis-specific p-value function \(p_T(\theta;x)\), the counternull set is schematically \(C_T(x,\theta_0)=\{\theta\ne\theta_0:p_T(\theta;x)=p_T(\theta_0;x)\}\). The notation does not guarantee that \(C_T\) is nonempty, unique, or independent of \(T\). A scalar counternull is a case where the construction identifies a single relevant nonnull member. The tests compared must be specified coherently: for opposite centers, “extreme” may point in opposite directions, as in the original normal illustration.[2]

What It Is Not

A counternull is not a confidence interval. An interval typically gathers parameter values under a coverage or test-inversion rule, whereas the counternull construction selects nonnull values lying at the same chosen test score as the original null. In a simple symmetric case it identifies another point, not an entire interval of equally supported values. A later randomization-based counternull Η set can contain multiple values, but set membership still follows equal-p matching, not ordinary interval coverage.[2]

It is not merely a doubled estimate. Doubling follows when the reference distribution and parameterization make the test symmetric around the estimate and the null is zero; with a nonzero null, the reflected value is \(2\hat\theta-\theta_0\) under that same geometry. Asymmetric models, discrete randomization distributions, nonlinear scales, or changed statistics require recomputation rather than automatic reflection.

It is not a posterior probability comparison, Bayes factor, guarantee of causal effect, clinical decision threshold, or claim that the null is false. Equal p-values report a property of the selected testing procedure. A large counternull may be scientifically interesting, but no importance threshold follows from the statistic itself. Conversely, a small counternull can focus attention on practical insignificance even where the original null is rejected.[1][2]

Scope of Application

The counternull arose in psychological effect-size reporting, especially as a corrective to reading a nonsignificant result as “no effect.” Rosenthal and Rubin also proposed its usefulness in meta-analysis, where a summary effect estimate can be accompanied by a nonnull comparison value.[1] The same construction can be stated for a trial treatment contrast, a regression coefficient, or another statistical estimand if a null, observed statistic and hypothesis-specific p-value calculation are defined. That is transfer within statistical inference, not a claim that counternulls independently appear in nonstatistical systems.

The 2025 randomization-test extension considers a sharp null hypothesis for a randomized experiment and calculates counternull sets under a chosen test statistic. Bind and Rubin's body-worn-camera study illustrates a consequential limit: different statistics applied to the same experiment can yield substantially different approximating counternull sets for one outcome.[2] Here the data, design, sharp hypotheses, statistic and numerical randomization procedure all matter. A scalar normal shortcut must not be carried into the exact randomization setting without checking these ingredients.

Clarity

To recognize a genuine counternull claim, ask five questions. What is the estimand and designated null value? What observed estimate or statistic is fixed? How does the selected procedure compute a p-value under each candidate value? Which nonnull value or values actually attain the null's score? Is the reported interpretation confined to that equality? An answer that gives only an effect estimate and a conventional null p-value lacks enough information to verify a counternull.

The seed's example \(d=0.25\) with p=0.20 omitted a standard error, sample size and test rule, so its numerical p cannot be checked. In the deliberately stipulated normal example below, we supply a fixed standard error and two-sided test before calculating anything. Likewise, “a 6-point risk reduction is as plausible as no reduction” would overstate what an equal-p calculation establishes. The defensible statement is narrower: under the specified test, those values get the same tail score.

Manages Complexity

Conventional reporting often compresses study results into a null p-value and a binary threshold verdict. The counternull adds one deliberately chosen counterpoint: a nonzero magnitude treated identically by the selected tail test. This can disrupt the false contrast between “zero” and “proved nonzero” without requiring every reader to inspect the full sampling distribution. The construction is especially communicative where a simple symmetric test makes the counterpart easy to calculate.

The compression has a cost. It hides the surrounding continuum of effect values, dependence on the selected test, and, in randomization settings, possible multiplicity of solutions. A responsible report keeps the original estimate, uncertainty analysis and test definition visible. The counternull supplements them; it should not replace a confidence interval, sensitivity analysis or substantive assessment of effect importance.[2]

Abstract Reasoning

For a symmetric two-sided normal test with fixed standard error \(s\), let the observed estimate be \(\hat\theta\) and the tested value be \(\theta\). The test distance is \(|\hat\theta-\theta|/s\). Equal two-sided tail areas arise at parameter values equally far from the estimate. Given null \(\theta_0\), the nonnull reflected counterpart is \(\theta_*=2\hat\theta-\theta_0\), provided it differs from \(\theta_0\). If \(\hat\theta=\theta_0\), the reflection lands on the null and no distinct scalar counternull emerges from this geometry. This is an algebraic explanation of the classical shortcut, not a claim about every test.[1][2]

The deeper operation is level-set comparison: hold observed data and scoring procedure fixed, then search for other hypotheses that yield the same score. A normal symmetric score can have a mirrored point; a discrete randomization score can have a plateau, multiple regions or no suitable nonnull match. Because the score itself depends on how extremeness is defined, the resulting set can move when the statistic or tail convention changes. The mechanism is not a test of which value is true; it is a way to expose what the specified test alone does and does not distinguish.[2]

Knowledge Transfer

To transfer the counternull from a psychology effect-size analysis to another empirical field, preserve the role structure rather than the numerical multiplier. Name the target estimand and its null, freeze the observed data, specify the model or randomization distribution and test statistic, and solve the equal-p relation for nonnull values. A result expressed as standardized mean difference cannot simply be transplanted to risk difference or odds ratio by multiplying a displayed number by two; symmetry must hold on the actual inferential scale.

In a randomized experiment, the role mapping can change further. One must state the sharp null family and statistic whose randomization distribution is being recomputed. The existence of a counternull Η set there is not evidence that every element is equally likely to be true. The cross-domain transferable insight is to ask what alternatives the same method treats like the null, while keeping the method's limitations attached to the answer.[2]

Examples

Illustrative psychology effect estimate. Suppose a standardized effect estimate is \(\hat d=0.25\), its stipulated standard error is $0.20$, and the analysis uses a two-sided normal test with the same standard error under candidate values. Testing \(d_0=0\) gives \(z=|0.25-0|/0.20=1.25\) and p approximately $0.211$. Testing \(d_*=0.50\) gives \(z=|0.25-0.50|/0.20=1.25\) and the same p. These numbers are a constructed illustration of Rosenthal and Rubin's symmetric case, not a report from an actual experiment. The equality does not make \(d=0.50\) clinically or psychologically consequential.[1]

Mapped back: The designated null is \(d=0\); the observed estimate is $0.25$ with stipulated \(s=0.20\); the test is two-sided normal with fixed \(s\); the nonnull candidate is \(d=0.50\); and the equal-p comparison is the shared absolute \(z\) distance of $1.25$.

Applied randomized body-worn-camera study. Bind and Rubin reanalyzed an experiment that randomized Washington, DC police officers to assignment to body-worn cameras. For the annual use-of-force rate, they tested a zero assignment effect against sharp constant-additive effects \(c\), using the reported assignment mechanism. With an adjusted regression treatment coefficient of 106.4 as the test statistic, 10,000 simulated allocations gave an approximate Fisher p near 0.94 under zero and an approximating counternull set \(c\in[216.963,216.982]\) in the study's annual-rate-per-1,000-officers units. Repeating the analysis with a Horvitz–Thompson statistic of 68.4 instead gave its own zero-null p near 0.84 and set \([141.031,141.056]\). The two different p-values are not claimed equal to each other: each nonnull set matches the zero-null score only within its own statistic and randomization test. Finite simulation and discrete randomization mean these are approximating sets, not exact continuous intervals of posterior plausibility. Noncompliance also means assignment effect is not identical to the effect of actually wearing cameras.[2]

Mapped back: The designated null is zero effect of camera assignment on use-of-force rate; the observed statistic is the adjusted regression coefficient 106.4 (or, in a separate analysis, Horvitz–Thompson 68.4); the test is a Fisher randomization calculation under each sharp constant-additive \(c\) and 10,000 simulated assignments; the nonnull candidate set is \([216.963,216.982]\) for regression (or \([141.031,141.056]\) for Horvitz–Thompson); and the equal-p comparison matches each candidate set to its own zero-null approximate Fisher p (about 0.94 or 0.84, respectively).[2]

Structural Tensions

Compact scalar communication versus design-sensitive exactness. The reflected point is a memorable communication aid when the test is symmetric, but its ease can tempt a reporter to apply it to a discrete or asymmetric test for which it is wrong. In a randomization analysis, different choices of statistic can produce different sets from the same experiment.[2] Diagnostic: Has symmetry, the tail rule and existence of a unique nonnull solution been checked under the actual analysis?

Correcting null acceptance versus overreading equal scores. Reporting a counternull can counter the inference “we did not reject zero, therefore the effect is zero.” Yet saying the two values are “equally plausible” without qualification can turn a p-value identity into an unwarranted posterior or likelihood claim. Diagnostic: Does the statement specify equality of p-values under a named test, and separately discuss practical importance and uncertainty?

Stable estimand versus changing scale. A reflected value on an additive effect scale need not stay reflected after a nonlinear transformation. This does not invalidate the counternull; it requires recomputing its score under the transformed hypothesis family. Diagnostic: Are null, estimate, candidate and uncertainty all expressed on the same inferential scale?

Structural–Framed Character

Evaluative weight: Low in the definition. The equal-p criterion is formal; judgments that an effect is useful, harmful or clinically large enter only after it is calculated. Human-practice dependence: Moderate. A researcher chooses the estimand, statistic, tail convention and reporting practice, but the calculation is constrained once those choices are fixed. Institutional origin: Low as a constitutive matter. Psychology and medical-research reporting norms motivate the statistic, but no particular journal, regulator or institution is required. Vocabulary travel: Moderate across empirical fields that use hypothesis tests; “counternull” is not an ordinary name for nonstatistical counterexamples. Import versus recognition: Mostly import. An analyst constructs the counternull under a declared test rather than discovering a natural object that exists independent of inference design.

The portable skeleton is “find a nondefault candidate on the same score level as a default candidate.” Whether such level-set comparison deserves a future prime is a separate identity question; the present entry does not admit that generalization as a prime or infer a cross-domain DAG parent from it. The source-established counternull remains bound to statistical testing and effect/parameter values. Its character: formal-statistical and structurally precise within a human-selected inferential frame, with interpretation requiring care.

Structural Core vs. Domain Accent

Structural core: Fix data, designated null and a hypothesis-specific scoring procedure; search for distinct candidates at the null's score level. In the sourced counternull, that score is a p-value. A point or set, rather than the numerical multiplier, is the stable identity.

Domain accent: Psychological standardized effects, treatment contrasts, zero nulls, normal approximations, randomization distributions, and decisions about clinical importance shape particular uses. None by itself defines the counternull. The twice-the-estimate slogan is a conditional accent of symmetric one-dimensional tests.[1][2]

Prime boundary: One could ask whether equal-score companion values constitute a portable higher-order abstraction outside statistics. That is an explicit future-prime question, not a finding here: these sources establish no independent nonstatistical cases or demonstrated substrate-general transfer.

This entry stages a strict composition/presupposes relation to Statistical Significance (p-Value): to define the counternull by equality, the analyst must be able to calculate the relevant p-value under each tested hypothesis. The counternull is not a kind of p-value and does not inherit a claim that equal p-values mean equal truth probability. Hypothesis Testing (Null vs. Alternative) supplies a broader setting; Effect Size supplies many quantities being compared; Confidence Intervals are a related but distinct uncertainty display. None is promoted as a duplicate or alias.

Relationships to Other Abstractions

Local relationship map for CounternullParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.CounternullDOMAINPrime abstraction: Statistical Significance (p-Value) — presupposesStatistical Sig…PRIME

Current abstraction Counternull Domain-specific

Parents (1) — more general patterns this builds on

  • Counternull presupposes Statistical Significance (p-Value) Prime

    A counternull is selected by matching the p-value calculated for a designated null under a specified test.

Hierarchy paths (11) — routes to 5 parentless roots

Neighborhood in Abstraction Space

Counternull sits in a sparse region of the domain-specific corpus (73rd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Statistical Hypothesis Tests & Diagnostics (9 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • An arbitrary alternative hypothesis: Any nonzero value may be an alternative; only a value meeting the declared equal-p condition is a counternull.
  • The observed estimate: In the symmetric scalar case the estimate is midway between null and counternull; it is normally not itself the counternull.
  • A confidence interval or an interval of equal plausibility: A set of matched p-values is defined differently from coverage, and the classical case yields a point, not an interval.
  • Equal posterior probability: A p-value is a tail probability conditional on a hypothesis and test model, not the probability that the hypothesis is true.[2]
  • Automatic practical importance: The counternull supplies a magnitude for discussion; substantive consequences require field-specific evidence and standards.[1]

References

[1] Robert Rosenthal and Donald B. Rubin, “The Counternull Value of an Effect Size: A New Statistic,” Psychological Science 5, no. 6 (1994): 329–334, especially publisher abstract defining equality of p-values and qualifying the twice-the-effect shortcut. Original publisher record. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i

[2] M.-A. C. Bind and Donald B. Rubin, “Counternull Sets in Randomized Experiments,” The American Statistician 79, no. 2 (2025): 275–285, especially Abstract, §§1–3, Figure 1 and randomized-study analysis. Original open-access paper. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p