Skip to content

Ecological Correlation

A correlation computed on aggregated group-level units that need not equal — and can reverse the sign of — the individual-level correlation, so it warrants no claim about the individuals inside the groups.

Core Idea

An ecological correlation is a correlation computed between variables measured at the group level — on aggregated units such as countries, counties, schools, or census tracts — rather than on the individuals within those groups. The critical fact, demonstrated formally by W. S. Robinson in 1950, is that an ecological correlation need not equal — and routinely does not equal, and can even have the opposite sign from — the corresponding individual-level correlation among the persons or cases being aggregated. Robinson's own example was stark: the 1930 U.S. census showed a cross-state correlation of +0.53 between the proportion of residents who were foreign-born and the proportion who were literate, but the individual-level correlation between being foreign-born and being literate was −0.11. Using the ecological (state-level) correlation to infer that foreign-born individuals were more likely to be literate would have inverted the true individual-level relationship.

The mechanism is that aggregation is a lossy operation: collapsing individual observations into group means discards within-group variance and leaves only between-group variance. When between-group variation in both variables is driven by a common third factor that also determines the group's composition — rather than by the individual-level relationship — the ecological correlation reflects that third factor and not the individual-level association. The direction and magnitude of the distortion depend on the partition chosen (which groups, how large, where their boundaries fall), the within-group distributions, and the degree to which group membership is itself confounded with other variables. Because the analyst controls only the group-level data, the distortion is invisible in the ecological statistic alone: no amount of inspection of the group-level numbers can reveal whether the individual-level relationship has the same sign, let alone magnitude. The practical implication is a firm inferential rule: an ecological correlation is not a warrant for any claim about individual-level associations; making such a claim from aggregate data is the ecological fallacy.

Structural Signature

Sig role-phrases:

  • the two variables on shared units — a pair of attributes measured on the same underlying individuals or cases
  • the aggregating partition — the chosen grouping (states, counties, tracts) that collapses individuals into groups, discarding within-group variance and keeping only between-group variance
  • the group-level coefficient — the correlation computed on the aggregated means: the ecological correlation itself
  • the level gap — the structural fact that this group-level coefficient need not equal, and can reverse the sign of, the individual-level correlation among the aggregated units (Robinson's +0.53 vs. −0.11)
  • the presumed-silence rule — the inferential warrant: a coefficient is silent about every level other than the one it was computed at until the within-group structure is modelled
  • the non-recoverability limitation — what the statistic deliberately cannot resolve: from the aggregate numbers alone the sign of the individual-level relationship is not recoverable, so the only remedy is individual-level data, not a covariate
  • the look-alike neighbours — the companion family that bounds the construct: sign-reversal (Simpson's paradox), partition-dependence (modifiable areal unit problem), and individual-rate recovery (the ecological inference problem)

What It Is Not

  • Not a warrant for any individual-level claim. Reading a group-level coefficient as if it described the persons inside the groups is precisely the ecological fallacy the construct exists to forbid. "States with more X have more Y" licenses nothing about "people with X have more Y"; an ecological correlation is presumed silent about every level other than the one it was computed at.
  • Not nothing about "ecology" or the environment. "Ecological" is a statistical term of art for the aggregate unit of analysis (Robinson's usage), not a reference to ecosystems or environmental variables. A correlation among countries' literacy rates is "ecological" regardless of subject matter; the word marks the grain, not the topic.
  • Not merely an attenuated or weakened version of the individual correlation. Aggregation does not just shrink the coefficient toward zero — it can leave it stronger, or reverse its sign entirely, as in Robinson's +0.53 against −0.11. The sign of the individual-level relationship is not recoverable from the aggregate numbers at all, so the ecological figure is not a noisy estimate of the individual one.
  • Not ordinary confounding. The distortion is aggregation loss, not an omitted covariate, so a control variable cannot fix it — only individual-level data recovers the relationship. The two problems can co-occur and must be diagnosed separately; reaching for a covariate when the trouble is grain is the wrong tool.
  • Not Simpson's paradox. Sign reversal on disaggregation is the dramatic special case; ecological correlation also covers non-reversing distortions that merely inflate, deflate, or otherwise misstate the relationship. Simpson's paradox is one way an ecological correlation can mislead, not a synonym for the phenomenon.
  • Not the ecological inference problem. Recovering individual-level rates from group marginals is a distinct constructive project; ecological correlation is the prior, prohibitive diagnosis that the group-level statistic is not the individual-level statistic. One tries to back out the micro-relationship; the other warns you not to assume the macro-coefficient already gave it to you.

Scope of Application

Because ecological correlation is a statistic-plus-inferential-rule, not a causal mechanism, it is not bounded by a single subject-matter domain: it applies wherever a correlation is computed on aggregated units and a claim is risked about the units inside them. The fields below are real uses of the identical construct — the same coefficient, the same Robinson-style sign-reversal hazard, the same presumed-silence default — not metaphor; the boundary is the construct's own precondition (a correlation, on aggregated units, with a different-level claim) versus over-reading it past that gap.

  • Epidemiology — country- or region-level exposure correlations (fat intake, sugar consumption) against disease prevalence, routinely misread onto persons; the canonical site of the warning.
  • Political science and sociology — precinct- or district-level returns read onto individual voters, where the closely related ecological inference problem lives.
  • Economics — cross-country GDP-per-capita correlations contrasted against the household-level relationships they do not establish.
  • Education research — school- or district-level test-score correlations used, illegitimately, as student-level proxies.
  • Public health and policy analysis — aggregate-data findings (county diabetes, regional mortality) that a control variable cannot rescue, the only remedy being individual-level data.

Clarity

Naming the ecological correlation forces apart two things a careless reading fuses: the unit of analysis at which a number was computed and the unit of claim an interpreter wants to make from it. The slide from "states with more X also have more Y" to "people with X have more Y" crosses a level of aggregation, and without a name for the group-level statistic that crossing is silent — the sentence reads as a single fact rather than as an inference smuggled across grain. Labelling the group-level coefficient as ecological, and the unlicensed leap as the ecological fallacy, puts the move in plain sight and converts an invisible error into a flagged one: the analyst can now point to exactly where the warrant was assumed rather than earned.

The label also supplies a working triage rule and the sharper questions behind it. Any correlation should carry the aggregation level at which it was measured, and should be presumed silent about every other level until the within-group structure is modelled — so the practitioner stops asking the unanswerable "is this relationship real?" and starts asking the locatable ones: at what grain was this computed, what partition produced the groups, and is the between-group variation that drives the coefficient the same relationship that holds among individuals or a different one imposed by group composition? The construct sharpens the boundary against neighbours that share its shape but are not it — Simpson's paradox as the special case where disaggregation reverses the sign, the modifiable areal unit problem as the dependence on how the groups were drawn, the ecological inference problem as the distinct project of recovering individual rates from group marginals. By distinguishing the aggregation distortion from ordinary confounding (a group-level correlation can suffer both at once), it tells the analyst that the remedy is not merely a control variable but individual-level data, because no inspection of the aggregate numbers alone can decide whether the underlying relationship even shares the sign the ecological figure displays.

Manages Complexity

Aggregate data are everywhere in social and health research — census tables, cross-country panels, precinct returns, school report cards — and any pair of group-level variables can be correlated, then read as if the coefficient spoke about the people inside the groups. Each such reading is, in principle, its own statistical investigation: one would have to model the within-group distributions, the partition, and the confounding structure to know whether the individual-level relationship matches. Ecological correlation compresses that open-ended verification problem to a single tracked attribute — the aggregation level at which the coefficient was computed — and a single default rule: a correlation is presumed silent about every level other than the one it was measured at, until the within-group structure is explicitly modelled. The analyst no longer asks the unanswerable "is this relationship real?" for each aggregate finding, but reads the qualitative status off one fact about provenance: computed on groups, it warrants nothing about individuals, full stop. A whole class of policy and journalistic claims — sugar-tax reasoning from country-level diabetes, voter inferences from precinct blocs — is pruned by that one move, before any modelling is attempted.

What makes the compression powerful is the structural fact that licenses the default: aggregation is lossy in a way no inspection of the aggregate can undo. Collapsing individuals to group means discards within-group variance and keeps only between-group variance, so when a common third factor drives both the between-group variation and the groups' composition, the ecological coefficient indexes that factor rather than the individual-level association — and can therefore differ in magnitude or even reverse in sign, as in Robinson's +0.53 against −0.11. Because the analyst holds only the group-level numbers, the sign of the underlying relationship is simply not recoverable from them, which is exactly why the blunt "assume silence" rule is the correct compression rather than a lazy one. The construct then organises the neighbourhood the analyst must read off the case: whether the distortion is a sign reversal on disaggregation (Simpson's paradox), a dependence on how the groups were drawn (the modifiable areal unit problem), the distinct task of recovering individual rates from group marginals (the ecological inference problem), or ordinary confounding that a control variable might address — versus the aggregation distortion proper, whose only remedy is individual-level data. So a sprawling "does this number mean what the writer says" problem reduces to tracking one parameter, the grain, applying one default, presumed silence, and placing the case in a small fixed branch set — from which the inferential verdict and the required remedy read off without re-deriving the statistics each time.

Abstract Reasoning

Ecological correlation licenses reasoning that polices the grain at which a statistic was computed against the grain at which a claim is made, and its moves are mostly prohibitive — they tell the analyst what cannot be inferred. The defining boundary-drawing move is the level rule: a correlation computed on groups warrants nothing about individuals, full stop, and is presumed silent about every level other than the one it was measured at until the within-group structure is explicitly modelled. Reasoning FROM "this coefficient was computed on aggregated units (states, counties, tracts)" TO "it licenses no claim about the persons inside those units" is what flags the slide from "states with more X have more Y" to "people with X have more Y" as the ecological fallacy — an inference smuggled across grain rather than a single fact.

The most distinctive feature is a negative diagnostic about recoverability: the analyst infers, before any modelling, that the sign of the individual-level relationship is not recoverable from the aggregate numbers at all. Because aggregation discards within-group variance and keeps only between-group variance, when a common third factor drives both the between-group variation and the groups' composition the ecological coefficient indexes that factor, not the individual association — so it can differ in magnitude or even reverse in sign, as in Robinson's +0.53 against −0.11. Reasoning FROM "I hold only group-level data" TO "no inspection of these numbers can tell me whether the underlying relationship even shares this sign" is what makes the blunt "assume silence" rule correct rather than lazy: the distortion is invisible in the ecological statistic by construction.

An interventionist move follows from that negativity and is unusually specific about the remedy. Because the trouble is aggregation loss, not ordinary confounding, the reasoner predicts that a control variable cannot fix it — the only remedy that recovers the individual-level relationship is individual-level data. Reasoning FROM "is this an aggregation distortion or a confound" TO "do I need micro-data or merely a covariate" is what keeps an analyst from reaching for the wrong tool, and it is why the construct insists the two diagnoses can co-occur and must be separated.

A classificatory boundary-drawing move places a given case among look-alikes that share the level-shift shape but differ in what to check. A sign reversal on disaggregation is Simpson's paradox; a dependence on how the groups were drawn is the modifiable areal unit problem; the constructive task of recovering individual rates from group marginals is the ecological inference problem; ordinary confounding is something a covariate might address. Reasoning FROM "does this case reverse, depend on the partition, ask for recovery, or merely confound" TO "which neighbor it is, and therefore which check applies" is what tells the analyst whether to inspect the partition, model the within-group distribution, or demand micro-data — rather than treating every aggregate anomaly identically.

Knowledge Transfer

Ecological correlation is not a causal mechanism but a statistic-plus-inferential-rule — a correlation coefficient computed at the group level together with the prohibition on reading it as an individual-level association — so the usual "mechanism within / metaphor beyond" framing does not apply to it. What transfers is the construct itself, literally, to any field that computes a correlation on aggregated units and then risks a claim about the units inside them. Within statistics and data science its precondition holds identically across substrates that differ only in subject matter: epidemiology (country-level fat-intake or sugar-consumption against disease prevalence, read onto persons), political science and sociology (precinct returns read onto individual voters — the home of the closely related ecological inference problem), economics (cross-country GDP-per-capita correlations versus household relationships), and education research (school-level test-score correlations used as student-level proxies). In every one the same coefficient, the same Robinson-style sign-reversal hazard (+0.53 against −0.11), the same presumed-silence default, and the same remedy (individual-level data, not a covariate) carry over without translation. This is instrument-reach, and it is wide: wherever there is grain to cross, the diagnostic applies verbatim.

The boundary to mark, then, is not analogy-versus-mechanism but instrument-reach versus over-reading, and it cuts two ways. First and most important, over-reading is the very error the construct names: the ecological fallacy is precisely reading the group-level statistic as if it spoke about individuals, and the entire value of the term is to forbid that over-read. Second, the construct should not be over-extended past its own precondition — it is silent unless there genuinely is (a) a correlation, (b) computed on aggregated units, with © a claim being made at a different level. Invoking "ecological correlation" loosely for any reasoning that mismatches scales, where no aggregated coefficient and no individual target are in play, borrows the name without the instrument and should be marked as such.

Where a genuinely cross-domain lesson is wanted — that summarising data at a coarser grain can hide, weaken, strengthen, or reverse the relation among the underlying items — that lesson is not unique to this construct and should be carried by the general primes it sits atop, not by "ecological correlation": aggregation (the lossy collapse that discards within-group variance), correlation (the relation being computed), and scaling_and_scale_dependence (relations changing qualitatively across grain) together entail the ecological-correlation phenomenon as a consequence. The named statistical furniture — Robinson's census example, the ecological/individual coefficient pair, the partition and within-group-distribution machinery, and the look-alike taxonomy (Simpson's paradox, the modifiable areal unit problem, the ecological inference problem) — is the working vocabulary of the statistical-research toolkit and stays home; the portable insight about scale-dependent relations belongs to the parent primes (see Structural Core vs. Domain Accent).

Examples

Canonical

W. S. Robinson's 1950 paper "Ecological Correlations and the Behavior of Individuals" supplies the defining worked instance. Using the 1930 U.S. Census, Robinson computed, across the 48 states, the correlation between the proportion of residents who were foreign-born and the proportion who were literate: r = +0.53. Taken at face value this suggests immigrants were more literate. But the correlation between being foreign-born and being literate computed on individuals was r = −0.11 — immigrants were, if anything, slightly less likely to be literate. The reversal arises because immigrants concentrated in northern states whose native-born populations were also more literate for unrelated reasons; the state-level coefficient indexes that geographic confounding of group composition, not the person-level relationship. The two numbers, +0.53 and −0.11, are opposite in sign.

Mapped back: Foreign-birth and literacy are the two variables on shared units (individuals); the 48 states are the aggregating partition. The +0.53 is the group-level coefficient, and its opposition to the individual −0.11 is exactly the level gap. That no reading of the state numbers alone reveals the person-level sign is the non-recoverability limitation, which is why the presumed-silence rule must hold.

Applied / In Practice

Émile Durkheim's Le Suicide (1897), a foundational work of empirical sociology, correlated suicide rates with religious composition across the regions and provinces of Europe, finding that predominantly Protestant areas had higher suicide rates than Catholic ones, and inferred that Protestants as individuals were more prone to suicide (which he tied to lower social integration). Later methodologists (notably Selvin, 1958) identified this as an ecological inference: the data are regional aggregates, and a higher regional suicide rate in Protestant-majority areas does not establish that it was the Protestants, rather than, say, Catholics living as minorities among them, who died. The substantive theory may hold, but the aggregate correlation alone cannot license the individual-level claim.

Mapped back: Regions are the aggregating partition; regional % Protestant and regional suicide rate are the two variables on shared units, and their association is the group-level coefficient. Durkheim's step to "Protestant individuals are more suicide-prone" crosses the level gap, and because individual-level data on who died were not in hand, the non-recoverability limitation applies — the case sits with the ecological-inference member of the look-alike neighbours.

Structural Tensions

T1: Presumed silence versus the atomistic fallacy in the mirror. The construct's core rule — an aggregate coefficient warrants nothing about individuals — is a vital guard against the ecological fallacy. But applied as a blanket "always demand micro-data," it tips into the opposite error, the atomistic (individualistic) fallacy: individual-level relationships can miss genuine contextual effects where the group is the operative causal unit (neighbourhood poverty acting on a resident beyond their own income, class composition acting on a pupil beyond their own ability). Sometimes the group level is not a lossy proxy for the individual level but the correct level of the claim. The tension is that the same discipline which stops an analyst reading persons off states can, over-applied, deny that groups ever exert effects individuals do not — reversing one fallacy into its twin. Diagnostic: Is the individual level genuinely the target here, or is the group itself the causal unit, so that demanding micro-data would erase a real contextual effect?

T2: The blunt non-recoverability claim versus the constructive ecological-inference project. The construct insists, pedagogically, that the sign of the individual relationship is "not recoverable from the aggregate numbers at all" — which is the right default and the reason "assume silence" is not lazy. Yet the neighbouring ecological-inference literature (Goodman, King) shows that partial recovery of individual rates from group marginals is possible under stated assumptions and auxiliary bounds. So the flat non-recoverability rule is a true safeguard and a slight overstatement at once: strictly, nothing is recoverable without assumptions, but something may be recoverable with them. The tension is that the prohibition's protective bluntness (never back out individuals) is in mild conflict with the fact that a disciplined constructive method can back out something, and treating recovery as flatly impossible forecloses a legitimate, caveated project. Diagnostic: Is the individual relationship being declared unrecoverable in principle, or merely unrecoverable without the explicit modelling assumptions ecological inference would supply?

T3: A reported coefficient versus the partition that produced it. An ecological correlation is presented as a number — Robinson's +0.53 — but that number is an artefact of a chosen partition: which units, how large, where the boundaries fall. Redraw the groups (the modifiable areal unit problem) and the coefficient shifts, sometimes dramatically, with no change in the underlying individuals at all. So the ecological correlation is not a fact about the data so much as a fact about the data under one grouping, yet it circulates with the authority of a single measured quantity. The tension is that reporting one coefficient conceals its dependence on an arbitrary aggregation choice — the same individuals can be made to yield many ecological correlations, and nothing in the reported number discloses which partition (or how gerrymanderable a one) generated it. Diagnostic: Would this coefficient survive a reasonable redrawing of the group boundaries, or is its value an artefact of the particular partition chosen?

T4: Aggregation as distortion versus aggregation as signal extraction. The mechanism is framed as loss — collapsing individuals to group means "discards within-group variance," which is what lets a third factor masquerade as the individual relationship. But the very same operation is how analysis averages out individual noise and makes stable group-level structure estimable in the first place; aggregation is not merely a corruption to be lamented but the tool that reveals relationships invisible at the individual grain. The operation that destroys individual-level warrant is the operation that manufactures group-level signal. The tension is that "aggregation is lossy" is true and one-sided: the loss of within-group variance that dooms individual inference is the same discarding of noise that makes the between-group relationship legible, so the construct's cautionary framing understates why anyone aggregates at all. Diagnostic: Is the aggregation here obscuring an individual relationship you actually want, or extracting a group-level regularity that the individual noise would have buried?

T5: Aggregation loss versus confounding (a boundary that is partly framing). The construct sharply distinguishes its distortion from ordinary confounding — "a control variable cannot fix it, only individual-level data can" — so the analyst reaches for micro-data rather than a covariate. Useful discipline. But the mechanism it describes is a confounding-by-composition story: the between-group variation is driven by a third factor entangled with group membership, which is confounding, and in some cases a group-composition covariate genuinely does attenuate the distortion. The clean "not confounding, need micro-data" boundary is therefore partly a matter of how the problem is framed rather than a hard type distinction. The tension is that insisting the only remedy is individual data can be wrong when the distortion is a modellable composition confound, while collapsing the two invites reaching for a covariate that cannot recover a sign the aggregate never contained. Diagnostic: Is the distortion here a compositional confound a group-level covariate could attenuate, or an aggregation loss that no covariate can undo and only micro-data can resolve?

T6: Autonomy versus reduction (a named statistical fallacy or the aggregation/scale-dependence parents). "Ecological correlation" is a precise statistical construct — a group-level coefficient plus the prohibition on reading it as individual-level — with home-bound furniture (Robinson's census pair, the partition machinery, the look-alike taxonomy of Simpson's paradox, MAUP, and the ecological inference problem), and within statistics it transfers literally to any field that correlates aggregates and risks an individual claim. But the portable lesson — that summarising at a coarser grain can hide, weaken, strengthen, or reverse the relation among the underlying items — is not unique to it and is entailed by the parents aggregation (the lossy collapse), correlation (the relation), and scaling_and_scale_dependence (relations changing across grain). The tension is that the named furniture is the working vocabulary of the statistical toolkit while the cross-domain insight belongs to those primes. Diagnostic: Resolve toward aggregation/scale-dependence when carrying the lesson about grain-changing relations generally; toward ecological correlation when a specific group-level coefficient is being wrongly read onto the individuals inside the groups.

Structural–Framed Character

Ecological correlation sits in the middle of the structural–framed spectrum — best read as mixed, because it is genuinely two things at once: a substrate-neutral mathematical phenomenon (the level gap) and a normatively-loaded, practice-bound inferential rule (the ecological fallacy), and the criteria split cleanly along that seam. On evaluative weight the phenomenon is neutral — a group-level coefficient is just a computed number, and Robinson's +0.53 against −0.11 is a fact about the data, praising and blaming nothing — but the construct's operative half is a prohibition: the ecological fallacy names a defective inference, a verdict that reading the aggregate onto individuals is illegitimate. That prohibitive core carries real evaluative weight in the way a fallacy label does, so the concept convicts an inferential move even though its underlying coefficient does not. On human-practice-bound the split is just as sharp: the mathematical phenomenon — that lossy aggregation can hide, weaken, strengthen, or reverse a relation — holds observer-free, true of the numbers whether or not any analyst computes them; but "ecological correlation" as a statistic-plus-inferential-rule, and the fallacy it forbids, are furniture of the practice of statistical inference, rules about what an analyst may claim, which have no referent absent that practice. Its institutional origin patterns the same way: the level gap is a mathematical consequence of aggregation, but the named construct, Robinson's 1950 coinage, and the look-alike taxonomy (Simpson's paradox, the modifiable areal unit problem, the ecological inference problem) are furniture of the statistical-research tradition — distinctions drawn inside the toolkit, not read off nature.

Where it diverges most from a domain-bound entry like isostasy is vocab_travels and import_vs_recognize, and here it is unusually mobile rather than pinned. Because it is a substrate-neutral instrument rather than a causal mechanism wearing a home vocabulary, the construct transfers literally — the same coefficient, the same sign-reversal hazard, the same presumed-silence default — to any field that correlates aggregates and risks an individual claim (epidemiology, political science, economics, education), so within quantitative work its reuse is recognition of the same instrument, not import-by-analogy. What stays home is not a substrate but the named statistical furniture — Robinson's census pair, the partition machinery, the look-alike taxonomy — and the prohibition's working vocabulary. The portable structural skeleton is a composite: summarising data at a coarser grain can hide, weaken, strengthen, or reverse the relation among the underlying items — the scale-dependence of a computed relation under lossy aggregation — and that skeleton is exactly what the catalog already carries as the parents aggregation (the lossy collapse that discards within-group variance), correlation (the relation being computed), and scaling_and_scale_dependence (relations changing qualitatively across grain), which together entail the ecological-correlation phenomenon as a consequence. More than one parent is genuinely load-bearing because the construct is precisely the intersection of a lossy aggregation, a correlation, and a scale-shift; the portable lesson belongs to those parents jointly, while the named fallacy and its census-and-taxonomy furniture stay in the statistical toolkit. Its character: a substrate-neutral scale-dependence phenomenon — evaluatively neutral and observer-free at its mathematical core — wrapped in a normatively-charged, practice-bound inferential prohibition and a tradition's named furniture, structural enough in its underlying aggregation/correlation/scaling_and_scale_dependence skeleton to travel as an instrument, framed enough in its fallacy-verdict and toolkit-specific vocabulary to remain mixed rather than a free-floating prime.

Structural Core vs. Domain Accent

This section pins down why ecological correlation is a domain-specific abstraction and not a prime, by separating the substrate-neutral scale-dependence phenomenon from the statistical-toolkit furniture and inferential prohibition that give it its identity — a case where the portable skeleton is genuinely composite.

What is skeletal (could lift toward a cross-domain prime). Strip the statistical vocabulary away and the surviving structure is a composite of three portable cores whose intersection is the phenomenon: a relation is computed between two variables (a correlation); the underlying items are collapsed to coarser units by a lossy summary that discards within-unit variation and keeps only between-unit variation (aggregation); and the computed relation can therefore change qualitatively — weakening, strengthening, or reversing — as the grain of measurement changes (scale-dependence). Each core is genuinely substrate-portable, which is exactly why the catalog already carries them as the parents the construct sits atop: correlation (the relation being computed), aggregation (the lossy collapse), and scaling_and_scale_dependence (relations changing across grain) — which together entail the ecological-correlation phenomenon as a consequence. This is the core the construct shares, held jointly by three parents, not what makes it distinctive.

What is domain-bound. What stays home is not a physical substrate but the named statistical furniture and the inferential rule the construct wraps around that skeleton: Robinson's 1930-census foreign-born/literacy coefficient pair (+0.53 against −0.11); the partition and within-group-distribution machinery; the look-alike taxonomy (Simpson's paradox as the sign-reversal special case, the modifiable areal unit problem as partition-dependence, the ecological inference problem as the constructive recovery project); the presumed-silence default and the non-recoverability claim; and above all the normatively-charged ecological fallacy — the prohibition that reading a group-level coefficient onto individuals is a defective inference. The decisive test: remove the correlation-on-aggregated-units-with-a-different-level-claim precondition and there is nothing for the fallacy verdict to attach to; invoking "ecological correlation" for any loose scale-mismatch, where no aggregated coefficient and no individual target are in play, borrows the name without the instrument. That worked vocabulary and its verdict are furniture of the statistical-research tradition — distinctions drawn inside the toolkit, not read off nature.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose transfer is recognition of the same mechanism, not analogy. Ecological correlation's transfer is bimodal, though its axis differs from a causal mechanism's: because it is an instrument rather than a substrate-bound process, within quantitative work it travels literally — the same coefficient, the same sign-reversal hazard, the same presumed-silence default, the same micro-data remedy carry verbatim across epidemiology, political science, economics, and education, because each supplies the construct's precondition; that is recognition of the same instrument, and its home "domain" is statistical inference itself. Beyond that precondition it travels only as a name — the fallacy verdict and the census-and-taxonomy furniture have no referent where there is no aggregated coefficient — so a bare appeal to "an ecological correlation" for any scale-mismatch is analogy, the boundary between the two. And when the genuinely cross-domain lesson is wanted — that summarising data at a coarser grain can hide, weaken, strengthen, or reverse the relation among the underlying items — it is already supplied, in more general form, by the parents the construct composes: aggregation, correlation, and scaling_and_scale_dependence, carried jointly. The cross-domain reach belongs to those parents; "ecological correlation," as named, carries statistical-toolkit baggage — Robinson's census pair, the partition and look-alike machinery, the fallacy prohibition — that stays in the toolkit and does not and should not travel as a prime.

Relationships to Other Abstractions

Local relationship map for Ecological CorrelationParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.EcologicalCorrelationDOMAINPrime abstraction: Correlation — is a kind ofCorrelationPRIMEPrime abstraction: Partition Dependence of Aggregates — is a kind ofPartition Depen…PRIMEDomain-specific abstraction: Fallacy of Division — presupposes, conditionalFallacy ofDivisionDOMAIN

Current abstraction Ecological Correlation Domain-specific

Parents (2) — more general patterns this builds on

  • Ecological Correlation is a kind of Correlation Prime

    Ecological correlation is a correlation specialized to variables measured on partition-aggregated groups rather than individuals.

  • Ecological Correlation is a kind of Partition Dependence of Aggregates Prime

    Ecological correlation is the correlation-coefficient species of partition dependence: grouping fixes the retained between-group variance.

Children (1) — more specific cases that build on this

  • Fallacy of Division Domain-specific presupposes, conditional Ecological Correlation

    In its ecological statistical branch, division presupposes a group-level correlation that the reasoner then reads onto individuals.

Hierarchy paths (2) — routes to 2 parentless roots

Not to Be Confused With

  • Simpson's paradox. The dramatic case where a relationship reverses sign when data are disaggregated (or aggregated) — a trend present in every subgroup vanishes or flips in the pooled data. Ecological correlation is the broader phenomenon: sign reversal is its most striking special case, but it also covers non-reversing distortions that merely inflate, deflate, or otherwise misstate the relationship without flipping it. Tell: did the association literally change sign across the aggregation levels (Simpson's paradox), or is the point the more general fact that the group-level and individual-level coefficients need not match (ecological correlation)?
  • Modifiable areal unit problem (MAUP). The finding that a group-level statistic depends on how the spatial (or other) units are drawn — their size and boundary placement — so redrawing the zones changes the coefficient with no change in the underlying individuals. Ecological correlation is about the level gap between group and individual; MAUP is about sensitivity to the partition itself. They co-occur (any ecological coefficient is partition-relative), but MAUP concerns the arbitrariness of the grouping, not the group-vs-individual leap. Tell: is the worry that a different carving of the groups would change the number (MAUP), or that the group number does not describe the people inside the groups (ecological correlation)?
  • Ecological inference problem. The constructive project of trying to recover individual-level rates or relationships from group marginals under stated assumptions (Goodman regression, King's method). Ecological correlation is the prior, prohibitive diagnosis that the aggregate coefficient is not the individual one; ecological inference is the attempt to back out the individual figure anyway, with caveats. Tell: are you forbidding the naive read of the aggregate (ecological correlation), or modelling to estimate the individual relationship from aggregates (ecological inference)?
  • Atomistic (individualistic) fallacy. The mirror error: inferring a group- or context-level relationship from individual-level data, thereby missing genuine contextual effects where the group is the operative causal unit. The ecological fallacy runs group→individual; the atomistic fallacy runs individual→group. Over-applying "always demand micro-data" tips into this twin. Tell: is the illegitimate leap from aggregates down to persons (ecological fallacy) or from persons up to contexts (atomistic fallacy)?
  • Ordinary confounding. A distortion caused by an omitted third variable correlated with both predictor and outcome — fixable, in principle, by measuring and controlling for that covariate. Ecological correlation's distortion is aggregation loss: a control variable cannot recover a sign the aggregate never contained, and only individual-level data can. The two can co-occur and must be diagnosed separately. Tell: could adding a covariate to the same-grain data fix it (confounding), or does the fix require descending to individual-level data because the grain itself is the problem (ecological correlation)?
  • Aggregation, correlation, and scale-dependence (the parent primes it composes). The substrate-neutral lesson — summarising at a coarser grain can hide, weaken, strengthen, or reverse the relation among the underlying items — is entailed jointly by aggregation (the lossy collapse), correlation (the relation), and scaling_and_scale_dependence (relations changing across grain). Ecological correlation is the statistical-toolkit instance with its census example, partition machinery, and fallacy verdict. Tell: for the general lesson about grain-changing relations, carry those parents; reserve "ecological correlation" for a specific group-level coefficient wrongly read onto the individuals inside the groups. (Treated fully in earlier sections.)

Neighborhood in Abstraction Space

Ecological Correlation sits in a sparse region of the domain-specific corpus (87th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (309 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12