Variance Decomposition Table¶
Analytic artifact — instantiates Ensemble and Population-Level Equilibrium versus Individual-Level Heterogeneity
Splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.
A Variance Decomposition Table takes the total spread that a stable aggregate conceals and partitions it into attributable sources — within-group, between-group, spatial, temporal, measurement, process. Its defining idea is that the useful question is rarely how much variation there is (any spread answers that) but what kind — because which source dominates tells you whether the heterogeneity beneath the equilibrium is harmless noise, real structure, or a signal that demands action. It is an offline analytic artifact: it explains a snapshot by attributing its variation, and it does not display live, fire alarms, or rule which claim is valid at which level.
Example¶
A semiconductor fab reports a steady average die yield of 92% across a product line — an equilibrium the finance team is content to plan around. Engineering wants to know what the 92% is hiding. A Variance Decomposition Table partitions the yield variation into named components: within-wafer (die-to-die across a single wafer), between-wafer, between-lot, between-tool (which etch chamber processed the lot), and measurement (probe repeatability). Ranked by share, the table is emphatic: the great majority of the spread is between-tool — one etch chamber runs about eight points below the rest of the fleet, dragging its lots down while the fleet average stays comfortable. The "stable 92%" turns out to be a blend of one sick tool and a healthy fleet. The outcome is that engineering pulls the offending chamber for maintenance rather than chasing a phantom process-wide defect, because the decomposition showed the heterogeneity was structural and localized, not random scatter.
How it works¶
- Choose the factors. Decide which sources the spread will be partitioned by — group, location, time, tool, measurement.
- Attribute the variance. Compute each factor's share of the total variation, using a nested or crossed variance-components analysis so overlapping factors are separated cleanly.
- Rank the sources. Order factors by contribution, so the dominant driver of the spread is unmistakable.
- Classify each source. Judge each contribution as noise, structure, or actionable signal — the step that turns a partition into a diagnosis.
Tuning parameters¶
- Factor set — which sources are separated. An omitted dominant factor gets misassigned to the ones you kept; too many factors starve each cell of data.
- Nested vs crossed structure — how factors relate (lots within tools, versus tools crossed with days). The wrong structure smears variance across the wrong sources.
- Measurement-error isolation — whether a gauge study peels off probe or instrument noise. Isolating it prevents real variation from being written off as measurement slop, and vice versa.
- Variance vs variance-share reporting — raw components or percentages of total. Shares communicate dominance; raw values preserve scale.
When it helps, and when it misleads¶
Its strength is that it converts an undifferentiated spread into an actionable diagnosis of source: it tells you not that the equilibrium hides variation but exactly what the variation is made of, so effort lands on the sick tool rather than the whole line.
Its failure mode is that it can only attribute variance to factors you chose to include, so an omitted driver is silently loaded onto the factors present, and thin cells produce false precision. The classic misuse is subgroup fishing — slicing by factor after factor until some "significant" source appears — which is exactly the discipline that analysis of variance was formalized to guard, provided the factors are named before the data are dredged.[n1] The guarding discipline is to predeclare the factor set, isolate measurement error with a gauge study, and treat any exploratory split as a hypothesis to confirm rather than a finding.
How it implements the components¶
microstate_variability_profile— its output is exactly a structured profile of the spread: total variation broken into components.heterogeneity_relevance_test— it classifies each partitioned source as noise, structure, or actionable signal.subgroup_and_locality_map— the between-group and spatial factors (tool, lot, wafer position) it partitions along.
It quantifies how much each source contributes but does not rule which claim is valid at which level, nor write the aggregation rule linking a member to the aggregate — level_of_analysis_boundary and aggregation_translation_rule — which is the work of Micro-Macro Crosswalk, its nearest twin. The split is clean: this table attributes the spread to sources (owning heterogeneity_relevance_test), while the crosswalk routes each claim to the level where it holds (owning level_of_analysis_boundary). It also does not monitor live — that is Distributional Dashboard.
Related¶
- Instantiates: Ensemble and Population-Level Equilibrium versus Individual-Level Heterogeneity — the anatomy of the variation an equilibrium conceals.
- Consumes: Stratified Sampling Review — a variance partition is only trustworthy on a representative sample, so it depends on the sample being certified first.
- Sibling mechanisms: Distributional Dashboard · Stratified Sampling Review · Micro-Macro Crosswalk · Agent-Based or Ensemble Simulation · Subgroup Excursion Alert · Representative Microcase Panel · Equilibrium Stress Test
Editorial Notes¶
Form Classification¶
Form family: Analysis, Modeling & Optimization
Rationale: Variance Decomposition Table operates as an analytical, modeling, inference, comparison, or optimization procedure that derives insight or a solution because it splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.
Independent corroboration: The frozen evidence defines Variance Decomposition Table as 'Splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is', so its operative form is Analysis, Modeling & Optimization.
Nearest alternative: Representation, Specification & Plan — Variance Decomposition Table includes features of a static representation, map, specification, schema, or prospective plan that externalizes information, but its defining operation is an analytical, modeling, inference, comparison, or optimization procedure that derives insight or a solution.
Review outcome: Independent reviewer agreement; medium confidence.
Origin Attribution¶
Primary origin: Statistics & Experimental Design
Origin pattern: Single lineage
Present-day reach: Universal
Rationale: Both independent reviews identify statistics experimental design as the historical home of the operation—Splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.. The retained alternates document formative adjacent traditions; the reach field, not the origin field, carries later applicability.
Related originating lineages:
- Accounting & Auditing — Accounting, auditing, and controlled-resource stewardship supplies a parallel or contributing lineage for the mechanism's defining operation: splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.
- Data Science & Analytics — Data science's modeling, validation, and monitoring tradition contributes a separate formative lineage to the mechanism's variance decomposition table logic.
- Mathematics — Mathematical modeling, proof, and abstract-structure practice supplies a parallel or contributing lineage for the mechanism's defining operation: splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.
Review resolution: Both blind reviewers independently place the defining operation—Splits the spread hidden beneath an equilibrium into named sources — within-group, between-group, temporal, measurement — so you can see what kind of heterogeneity it is.—in statistics experimental design. Their queued differences are secondary: alternate_origin_disagreement, origin_mode_disagreement, domain_reach_disagreement, encyclopedia_synthesis_disagreement. Reviewer A uniquely contributes no additional alternate; reviewer B uniquely contributes ['accounting_auditing', 'mathematics']. I preserve the full evidence-supported union of 3 alternate domain(s), without a numeric cap. origin_mode=single_lineage reflects the more specific lineage judgment in reviewer B's evidence, while domain_reach=universal separately records present-day portability. The affirmative encyclopedia-synthesis finding is preserved, and confidence=high uses the more conservative reviewer level.
Encyclopedia synthesis: The exact catalogued form synthesizes established practice rather than reproducing a single standard historical label.
Review outcome: Reconciled after independent review; high confidence.
Notes¶
[n1] Analysis of variance — R. A. Fisher's method for partitioning the total variation in a dataset into components attributable to distinct sources (treatments, blocks, error). Variance decomposition generalizes it to nested and crossed factors; its validity still rests on the factors being specified before the data are searched for effects. ↩