Stratified Residual Review¶
Disaggregation review — instantiates False Convergence Prevention
Breaks a stable aggregate into subgroups, residuals, and edge cases to expose the pockets where the system has not actually converged even though the average looks settled.
An average can be perfectly stable while the thing underneath it is anything but. Stratified Residual Review takes a settled aggregate — a headline metric that has flattened, a mean that has stopped moving — and takes it apart, splitting it into subgroups, residuals, and edge cases to find the pockets where convergence has not actually happened. Its defining idea is that a summary signal can hide unresolved variation: the aggregate looks converged because opposing movements in different strata cancel, or because a small hard-hit group is drowned out by a large comfortable one. The review's whole job is decomposition — it does not disturb the system or reproduce it, it reads the same data at a finer grain and maps where the residual variation concentrates, so a mean that hides a harmed subgroup cannot pass as genuine stability.
Example¶
A national childhood-vaccination program reports that aggregate coverage has reached its 90% target and held there for three consecutive quarters. On the dashboard the line is flat and above the goal; the program office is ready to call the campaign converged and redirect funds elsewhere. But the 90% is a single national number, and a national number is exactly the kind of summary that can be stable on top and unsettled underneath.
Stratified Residual Review disaggregates it. Split by region and household income, the flat aggregate comes apart: dense urban districts sit near 98%, while rural low-income districts languish below 65% — and because the urban population is far larger, its high coverage mathematically buries the rural shortfall in the national mean.[n1] The "converged" 90% was hiding a subgroup that never converged at all, one where an outbreak is most likely to start. The review produces a map of where the residual variation lives — which districts, which income bands, which age cohorts — turning a comfortable aggregate into a specific, addressable gap while the campaign is still funded to close it.
How it works¶
- Choose the stratification dimensions. Decide how to cut the aggregate — by subgroup, region, time window, residual, or edge case — guided by where hidden variation would do the most damage if it existed.
- Disaggregate the metric. Recompute the settled signal within each stratum instead of over the whole, so opposing or drowned-out movements stop cancelling.
- Compare each stratum to the standard. Hold every slice against the same convergence bar the aggregate claimed to meet, and flag the ones that fall short.
- Map the residual variation. Show where the shortfall concentrates and how large it is, so a diffuse worry becomes a specific list of non-converged pockets that a gate or reopening can act on.
Its distinguishing move is that it works within the existing sample, reading the same evidence more finely — not perturbing the system, and not fetching a fresh independent sample.
Tuning parameters¶
- Stratification granularity — how finely the aggregate is cut. Fine strata surface small hidden pockets but shrink each group toward noise; coarse strata are stable but can re-hide the very variation being hunted.
- Dimension choice — which cuts are examined. The harmful non-convergence hides along some dimension, and a review that only slices the convenient ones will miss the cut where the problem actually lives.
- Residual threshold — how large a subgroup deviation must be to count as a flagged pocket rather than expected spread. A low threshold catches real gaps but raises false alarms; a high one is clean but can wave through a genuinely stranded group.
- Edge-case emphasis — how much weight goes to the tails versus the mass. Heavy tail emphasis catches rare but serious failures; heavy mass emphasis keeps the review representative but can ignore the worst-off few.
When it helps, and when it misleads¶
Its strength is catching the harm an average conceals — the subgroup left behind, the region that never caught up, the edge case the summary statistic silently absorbed — while the process still has room to revise. It is the direct answer to the archetype's warning that aggregate signals hide subgroup differences, and it works on data already in hand, so it is often the cheapest probe available.
Its failure mode is the multiple-comparisons trap in reverse of its virtue: slice finely enough and some subgroup will always look anomalous by chance, so an over-eager review manufactures false pockets and cries wolf, while a lazy one that slices only the obvious dimensions misses the harmful cut entirely. Small strata compound the problem — a subgroup of a dozen cases is noise dressed as signal. The classic misuse is the inverse of the mechanism: reporting only the aggregate and never disaggregating at all, which is the exact failure this review exists to prevent. The guarding discipline is to pre-specify the strata and the residual threshold before mining the data, and to guard subgroup claims against multiple-comparison false positives so a real stranded group is not lost among statistical mirages.
How it implements the components¶
Stratified Residual Review realizes the decomposition side of the archetype — the parts that pull a settled aggregate apart to surface variation it conceals:
hidden_variation_probe— it decomposes a summary metric into strata to search out the unresolved differences the aggregate hides, exactly the probe's remit of testing edge cases and comparing contexts a headline number smooths over.residual_variation_map— it produces the map of where residual variation concentrates across subgroups and residuals, naming the non-converged pockets rather than letting them stay averaged away.
It does not inject a disturbance into the system (perturbation_test — Perturbation Probe), reproduce the result through an independent actor or fresh sample (independent_check, out_of_sample_check — Independent Replication), or supply the after-closure route to revisit a decision (review_or_appeal_path, reopening_rule — Appeal or Reopening Review). This review reads the existing data at a finer grain; it neither shocks the system nor re-runs it.
Related¶
- Instantiates: False Convergence Prevention — this is the mechanism that decomposes a stable aggregate to reveal the pockets where convergence has not actually occurred.
- Sibling mechanisms: Perturbation Probe · Sensitivity Testing · Independent Replication · Appeal or Reopening Review · Assumption Audit
Editorial Notes¶
Form Classification¶
Form family: Assessment, Review & Assurance
Rationale: Stratified Residual Review operates as a bounded evaluation of existing evidence or work that produces a finding or disposition because it breaks a stable aggregate into subgroups, residuals, and edge cases to expose the pockets where the system has not actually converged even though the average looks settled.
Independent corroboration: The frozen evidence defines Stratified Residual Review as 'Breaks a stable aggregate into subgroups, residuals, and edge cases to expose the pockets where the system has not actually converged even though the average looks settled', so its operative form is Assessment, Review & Assurance.
Nearest alternative: Representation, Specification & Plan — Stratified Residual Review includes features of a static representation, map, specification, schema, or prospective plan that externalizes information, but its defining operation is a bounded evaluation of existing evidence or work that produces a finding or disposition.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Convergent development
Present-day reach: Multi-domain
Rationale: Reviewing subgroup residual pockets tests convergence beyond averages.
Related originating lineages:
- Data Science & Analytics — Residual analytics find local error.
- Engineering & Design — Edge cases threaten performance.
- Mathematics — Mathematical modeling, proof, and abstract-structure practice supplies a parallel or contributing lineage for the mechanism's defining operation: breaks a stable aggregate into subgroups, residuals, and edge cases to expose the pockets where the system has not actually converged even though the average looks settled.
Review resolution: The blind reviewers agree that statistics_experimental_design is the primary origin and differ only on alternate origin disagreement. I preserve every independently explained alternate from both records rather than imposing a numeric cap. I retain convergent because the combined evidence shows independent disciplinary development. The broader reach of multi_domain records portability separately from historical provenance; encyclopedia_synthesis=true preserves the affirmative synthesis judgment where either reviewer identified one.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Reconciled after independent review; medium confidence.
Notes¶
[n1] Simpson's paradox is the phenomenon in which a trend or level that holds in aggregate reverses or disappears once the data are split into subgroups — a stable, on-target overall figure can conceal a subgroup moving the opposite way. It is the formal reason a converged average is not evidence that every stratum has converged, and why disaggregation is a distinct check rather than a redundant one. ↩