Tensions in Practice: Separate claim families in tension with one combined error budget¶
Four invented hypothesis-test results
A p-value measures how extreme a test result is under its null model; it is not the probability that the null is true. Here four invented valid p-values stay fixed. A declared rule divides a 0.05 error budget by the number of tests in each family. Two families of two use cutoff 0.025; one family of four uses 0.0125. H2 changes decision because the scope of the error promise changes.
Keep claims separately scoped
Give two distinct predeclared claim families their own error budgets.
Control the combined family
Bound the chance of any false rejection across all four claims.
Why these aims pull against each other
Broader family control lowers the per-test cutoff; separately scoped control does not inherit a combined 0.05 guarantee.
Choose an arrangement to see what changes and what remains difficult.
Rows keep the same four tests and p-values. The family boundary changes the cutoff and H2’s decision; “Not rejected” does not establish the null. Shading isolates H2’s changed decision.
What this choice protects
What it costs
When it fits
Compare the arrangements
Two declared families
Treat H1–H2 and H3–H4 as distinct families, each with budget 0.05. Reject the null exactly when p is at or below the cutoff.
| p | Cutoff | Decision | |
|---|---|---|---|
| H1 | 0.01 | .025 | Reject |
| H2 | 0.02 | .025 | Reject |
| H3 | 0.03 | .025 | Not rejected |
| H4 | 0.04 | .025 | Not rejected |
- What it protects
- H2 meets the less stringent within-family cutoff.
- What it costs
- The union of both families is not promised a 0.05 error bound.
- When it fits
- Fits genuinely separate, prospectively declared reporting targets with that limited scope clearly stated.
Illustration note: Reject means reject the null by this rule; not rejected does not mean proven true.
One combined family
Apply one 0.05 budget to H1–H4. Reject the null exactly when p is at or below the cutoff.
| p | Cutoff | Decision | |
|---|---|---|---|
| H1 | 0.01 | .0125 | Reject |
| H2 | 0.02 | .0125 | Not rejected |
| H3 | 0.03 | .0125 | Not rejected |
| H4 | 0.04 | .0125 | Not rejected |
- What it protects
- The declared error promise covers all four claims together.
- What it costs
- H2 is no longer rejected, potentially sacrificing detection of a real effect.
- When it fits
- Fits when any reported finding in this combined collection is part of the same claim family.
Illustration note: The cutoff rule assumes valid individual p-values; its union bound does not require independent tests.
What this illustration does—and does not—establish
The source supplies the tension. The invented setting, alternatives and any numbers illustrate a limited comparison; each arrangement retains its stated costs and conditions.
- All p-values are invented; no empirical effects or study quality are inferred.
- Family membership must not be changed after seeing which boundary yields a desired result.
- This is the elementary budget-per-test rule, not a comparison of false discovery rate procedures, replication or causal validity.
Source entries
Multiple Comparisons Correction
This source passage supplies the contextual tension. The concrete arrangements and schematic examples are editorial illustrations, not measured findings.
Family definition as researcher judgment call
The boundary of "the family" of tests is a researcher choice with no uniquely correct answer. Is the family a single paper (10 tests)? A research program (100 tests across 5 papers)? A field's entire literature (10,000 tests over decades)? The larger the family boundary, the more aggressive the correction required, the lower the resulting power.
The source operation
(2) Multiple comparisons correction is the family of statistical techniques that adjust either the per-test significance thresholds or the p-values themselves to control some specified error rate at the family-wise level — the family-wise error rate (FWER, probability of any false rejection), the false discovery rate (FDR, expected proportion of false discoveries among rejections), or alternative criteria like per-family error rate or false coverage rate.