Equity premium puzzle¶
Confront one number against one model — the ~6-point historical equity premium against what a consumption-CAPM with plausible risk aversion can rationalize — and read the order-of-magnitude miss as indicting a load-bearing assumption in an enumerable stack.
Core Idea¶
The equity premium puzzle is the observation, formalized by Rajnish Mehra and Edward Prescott in 1985, that the historical excess return of US equities over short-term government bonds — averaging roughly 6 percentage points per year over the 1889–1978 period — is orders of magnitude larger than what a standard consumption-based asset-pricing model can rationalize given plausible values of risk aversion. The model at issue is the representative-agent consumption CAPM with constant relative risk aversion (CRRA) preferences: in this framework, the equity premium compensates investors for the covariance of equity returns with aggregate consumption growth. Because consumption growth in the postwar US has been both low in volatility and weakly correlated with stock returns, the model requires implausibly high risk aversion — roughly 30–40 on some calibrations, versus the commonly assumed range of 1–10 — to match the observed premium. The puzzle is therefore not about why equities have high returns, but about why the gap between equity and bond returns is so large relative to the smoothness and modest risk of observed consumption fluctuations: either agents are extraordinarily risk-averse by any behavioral standard, or the consumption data misrepresents the risk agents actually face, or the expected-utility framework is wrong. A constellation of proposed resolutions has developed around each of those three branches: preference-side alternatives (habit formation, Epstein-Zin recursive utility separating risk aversion from intertemporal substitution, loss aversion, ambiguity aversion); rare-disaster and tail-risk arguments that average consumption is smooth but the true distribution includes rare catastrophic events whose probability is underweighted in postwar samples; limited-participation and market-friction arguments that the marginal investor in equity markets is not the representative consumer but a wealthier, less constrained subgroup whose consumption is more volatile and equity-correlated; and survivorship bias in the choice of the US stock market as the reference data series. None of the proposed resolutions has achieved consensus; the puzzle remains the central unsolved calibration anomaly in asset pricing after four decades of active research.
Structural Signature¶
Sig role-phrases:
- the calibrated model — the representative-agent consumption-CAPM with CRRA preferences, which prices the equity premium as the price of the covariance of equity returns with aggregate consumption growth
- the observed quantity — the historical equity premium, the ~6-percentage-point annual excess return of US equities over short-term government bonds
- the independently constrained parameter — the risk-aversion coefficient, pinned by behavioral evidence to roughly 1–10
- the order-of-magnitude gap — the model reproduces the premium only at risk aversion of 30–40, an order-of-magnitude miss too large to be a tuning failure
- the enumerable assumption stack — the short list of load-bearing premises: time-separable expected utility, CRRA preferences, smooth weakly-correlated consumption, a representative agent who is the marginal investor, an unbiased return sample
- the four-way branch of resolutions — every candidate fix must abandon exactly one layer: preference-side (habit, Epstein-Zin, loss aversion), tail-risk (rare disasters), limited-participation (marginal investor not representative), or survivorship bias (the data series itself)
- the differential fingerprint — each branch implies a distinct testable signature (disaster-economy premia, disaggregated stockholder consumption, lower foreign-market premia), so the puzzle says where to look to confirm or refute each
- the persistence (not arbitrage) caveat — the spread persists and compensates real (if mismeasured) risk rather than being riskless profit, blocking the "large spread, therefore free money" inference
What It Is Not¶
- Not the observation that equities earn high returns. The puzzle is the gap between the historical premium and what a consumption-CAPM with plausible risk aversion can rationalize — an order-of-magnitude miss — not the level of equity returns. Mehra and Prescott's sharpened framing even relocates the anomaly to the other leg: why is the risk-free return so low that mildly risk-averse agents won't bid it up?
- Not an arbitrage or free money. The premium persists and compensates real (if mismeasured) risk; it is not a riskless profit that trading would erase. Reading "a large, durable spread" as an exploitable mispricing inverts the puzzle, whose whole difficulty is that the spread survives precisely because it is not arbitrage.
- Not explained by "stocks are riskier than bonds." That intuition is exactly what fails. Written as the price of equity's covariance with consumption growth under CRRA preferences, the risk story generates only a fraction of a percentage point of premium at any defensible risk-aversion level. The qualitative "equities are risky" answer is not in dispute; the puzzle is that it cannot reproduce the magnitude.
- Not a verdict that markets are irrational or inefficient. The anomaly indicts the model — the representative-agent consumption-CAPM and its stack of assumptions (time-separable expected utility, the postwar consumption sample, the representative investor as marginal investor, an unbiased return series) — one of which must be misspecified. It is a calibration failure of a particular framework, not evidence that investors behave irrationally.
- Not a general cross-domain pattern that travels under its name. The equity premium puzzle is one asset-pricing anomaly, scoped to a specific model and a specific return series; there is no second domain in which "the equity premium puzzle" recurs as the same object. What is portable is the parent — a calibrated model diverging from a measured quantity by more than tuning can absorb, indicting an assumption, an input, or the frame — not this named puzzle, whose consumption-covariance machinery is asset-pricing furniture.
Scope of Application¶
The equity premium puzzle is a single calibration anomaly scoped to asset pricing, not a mechanism that recurs across subfields; its reach is depth on one number within a small cluster of consumption-CAPM anomalies, and the general "calibrated model vs. measured outcome" reasoning it instantiates travels cross-domain only under the parent calibration-anomaly pattern, not under this name. The habitats below are the within-domain contexts where its apparatus operates.
- Consumption-based asset pricing — the home turf, where the puzzle is the central unsolved calibration target: the ~6-point premium a representative-agent consumption-CAPM with CRRA preferences cannot reproduce at defensible risk aversion.
- The risk-free-rate puzzle — its mirror image (why is the bond return so low that mildly risk-averse agents won't bid it up), the same anomaly seen from the other leg.
- Preference-theory research — the branch that resolves the puzzle by abandoning time-separable expected utility: habit formation, Epstein-Zin recursive utility, loss aversion, ambiguity aversion.
- Rare-disaster / tail-risk macro-finance — the branch arguing the postwar consumption sample omits catastrophic-event probability, tested against disaster-economy premia and option-implied tail risk.
- Limited-participation and market-friction asset pricing — the branch denying the representative consumer is the marginal investor, pointing at disaggregated wealthy-stockholder consumption.
- Long-run equity-return measurement — the survivorship-bias branch, examining whether the US series is an upward-biased reference and whether interrupted foreign markets show lower premia.
Clarity¶
The puzzle's clarifying work is to relocate the question. The naive reading of high stock returns is "equities are risky, so they pay more" — a story that feels complete and explains nothing quantitatively. By writing the premium as the price of one specific covariance — equity returns with aggregate consumption growth — under CRRA preferences, Mehra and Prescott convert a vague intuition into a calibration: a single number the model must reproduce from independently constrained inputs. Once stated that way, the disturbing fact becomes visible. The premium is not merely large; it is larger than the consumption-based model can generate by an order of magnitude at any defensible risk-aversion level. Naming the puzzle is what makes that gap a fact to be explained rather than a slogan to be repeated, and it reframes the open problem precisely: not "why are equities risky?" but "why is the bond return so low — why won't agents who appear only mildly risk-averse bid it up?"
Its sharper service is to make every proposed fix declare which load-bearing assumption it abandons. The model rests on a small stack — time-separable expected utility, CRRA preferences, smooth and weakly-correlated consumption, a representative agent who is the marginal investor, an unbiased return sample. The puzzle's framing forces any resolution to name its target: habit formation and Epstein-Zin attack the preference structure; rare-disaster arguments attack the claim that the postwar consumption sample captures the true tail; limited-participation arguments deny that the representative consumer is the marginal investor; survivorship-bias arguments impugn the data series itself. An asset-pricing economist can therefore read a candidate resolution as a wager on exactly one assumption, compare rival resolutions on common ground, and ask the diagnostic question the puzzle exists to sharpen — which element of the consumption-CAPM is misspecified — rather than defending or attacking the framework wholesale.
Manages Complexity¶
The raw material an asset-pricing economist would otherwise have to confront is enormous: ninety-plus years of returns on hundreds of securities, the full joint distribution of equity and bond yields against consumption, and an open-ended list of candidate stories for why stocks pay more. The puzzle compresses all of it to a single confrontation between one number and one model. On one side, the historical equity premium — roughly six percentage points per year. On the other, what the representative-agent consumption-CAPM with CRRA preferences predicts that premium should be, given an independently constrained risk-aversion parameter. The entire returns distribution, with its dozens of moments and securities, is collapsed to the question "can the model reproduce this one spread at a defensible risk-aversion level?" — and the answer, an order-of-magnitude miss, is what the economist tracks instead of the data sprawl. Decades of evidence reduce to a single calibration target the model fails to hit.
That compression does more than shrink the data; it imposes a fixed branch structure on the otherwise unbounded space of explanations. Because the failing model rests on a short, enumerable stack of assumptions — time-separable expected utility, CRRA preferences, smooth and weakly-correlated consumption, a representative agent who is the marginal investor, an unbiased return sample — every conceivable resolution must abandon at least one named element of that stack, and the puzzle sorts the whole literature into a handful of mutually exclusive bets accordingly: attack the preferences (habit formation, Epstein-Zin, loss aversion, ambiguity aversion), attack the consumption sample's tail (rare disasters), attack the identity of the marginal investor (limited participation), or attack the data series itself (survivorship bias). An economist meeting a new proposed fix does not have to evaluate it against the full apparatus of asset pricing; the puzzle's framing lets them read it as a wager on exactly one branch, place it beside its rivals on common ground, and ask the single diagnostic question the whole structure exists to sharpen — which element of the consumption-CAPM is misspecified. A field-spanning controversy is thereby reduced to one anomaly, one short list of load-bearing assumptions, and a four-way branch, so that the qualitative shape of any candidate answer reads off from which assumption it sacrifices rather than from re-litigating the framework whole.
Abstract Reasoning¶
The equity premium puzzle licenses reasoning moves that are characteristic of how asset-pricing economists work with a calibration anomaly — moves that run from a measured gap to a misspecified assumption, and from a candidate fix to a falsifiable prediction.
The central move is diagnostic backward inference from the size of the miss. The premium is written as the price of one covariance — equity returns against aggregate consumption growth — under CRRA preferences, which means the model's implied premium is a function of the risk-aversion parameter and the observed consumption moments. Confronting a six-point premium that the model can generate only at risk aversion of 30–40, the economist reasons backward: since the consumption moments are independently measured and the risk-aversion range 1–10 is independently constrained by behavior, an order-of-magnitude miss cannot be a tuning failure within the model but must indict one of the inputs to the model. The magnitude of the gap is itself the diagnostic — a small miss would invite recalibration, but a factor-of-ten miss licenses the inference that a load-bearing assumption is false rather than merely mis-set. This is the move that converts "stocks pay more because they're risky" into a quantitative contradiction demanding structural explanation.
A second move is assumption-localization, treating the failing model as a short enumerable stack and reading any proposed resolution as a wager on exactly one layer. The economist holds the stack explicit — time-separable expected utility, CRRA preferences, smooth weakly-correlated consumption, a representative agent who is the marginal investor, an unbiased return sample — and asks of every candidate fix: which assumption does this abandon? Habit formation and Epstein-Zin recursive utility are read as bets that the preference structure is wrong; rare-disaster arguments as bets that the postwar consumption sample misses the true tail; limited-participation arguments as bets that the marginal investor is not the representative consumer; survivorship-bias arguments as bets that the data series itself is biased upward. The discipline of this move is that it puts rival resolutions on common ground — each is comparable as a claim about one named layer — and forbids the lazy response of attacking or defending "the model" wholesale.
A third move is interventionist in the sense of generating differential predictions: because each branch localizes the misspecification differently, each implies a distinct empirical signature, and the economist reasons from the branch to where it should leave a fingerprint. If the resolution is rare disasters, then the true return distribution should carry a fat left tail that postwar US data underweights, so the prediction is that economies which actually suffered catastrophes show premia consistent with the model, and that option prices reveal the disaster probability the sample omits. If the resolution is limited participation, then the consumption of the wealthy stockholding subgroup — not aggregate consumption — should be volatile and equity-correlated enough to rationalize the premium, so the prediction points at disaggregated stockholder-consumption data. If the resolution is survivorship bias, then other national equity markets, including those that were interrupted or expropriated, should show systematically lower premia. The puzzle thus does not merely catalog stories; it tells the economist what each story commits to and where to look to confirm or refute it.
A fourth move is a boundary-drawing one that keeps the puzzle from being misclassified as something it is not. The economist must distinguish the equity premium puzzle from an arbitrage: an arbitrage is a riskless profit that trading away would erase, whereas the premium persists and compensates real (if mismeasured) risk, so the move is to refuse the inference "large persistent spread, therefore free money" and instead read the spread as evidence of unmodeled risk or misspecified preferences. The puzzle also locates itself by reframing which leg is anomalous: the sharpened question is often not "why is the equity return high?" but "why is the bond return so low — why won't seemingly mildly-risk-averse agents bid it up?" — a relocation that points resolutions toward the risk-free rate puzzle as the same anomaly seen from the other side.
Knowledge Transfer¶
Within asset pricing the puzzle transfers as a live research object, and what carries is the full apparatus: the single calibration target (one number against one model), the short enumerable assumption stack, the four-way branch of resolutions, and the differential-prediction reasoning that tells each branch where to leave a fingerprint. But the honest within-domain characterization is that the EPP is one anomaly, not a pattern that recurs across the subfields — it does not "transfer" the way a mechanism does, because there is no second asset-pricing setting in which the same gap between a calibrated consumption-CAPM and observed risk-asset returns reappears as a distinct instance. What it does instead is anchor a four-decade program and sit inside a small cluster of related consumption-CAPM anomalies (the risk-free-rate puzzle as its mirror image, plus the volatility and value-premium puzzles), which together point at the same over-narrow utility framework. Its reach within the domain is depth on one calibration, not breadth across many.
Beyond asset pricing the report is twofold. (1) Moving the name outside its home strips both the model and the data: a "compensation premium puzzle" in labor or any other "X-premium puzzle" borrows the problem-statement form while having no consumption-CAPM, no calibrated risk-aversion parameter, and no equity return series, so it is analogy and should be marked as such. (2) The genuinely portable content is one level up, and it is a shared abstract pattern rather than a metaphor: the EPP instantiates the general structure a calibrated model diverges from a measured quantity by a factor too large to hand-wave → indict a load-bearing assumption, a mismeasured input, or the wrong frame, together with the reasoning that pattern licenses — backward inference from the size of the miss, localization of the failure to one named layer of an explicit assumption stack, and derivation of differential empirical signatures from each candidate resolution. That anomaly-resolution structure recurs across the empirical sciences as co-instances (any field where an independently-constrained model misses an independently-measured target by orders of magnitude reasons the same way), and it is the level at which the cross-domain lesson lives. The discipline to keep is that this is the parent (the side-captured calibration-anomaly / model-data-gap pattern), not the equity premium puzzle: the cross-domain reach belongs to "calibrated model vs. measured outcome, diverging beyond tuning," while the EPP's specific cargo — the consumption covariance, CRRA preferences, the 6-point premium, the rare-disaster/limited-participation/survivorship branches — is asset-pricing furniture that does not and should not travel. A live calibration anomaly within asset pricing; a shared abstract reasoning pattern — carried by the parent, not this named puzzle — beyond. This is exactly the boundary Structural Core vs. Domain Accent draws.
Examples¶
Canonical¶
The defining instance is Mehra and Prescott's own 1985 calibration in "The Equity Premium: A Puzzle." They measured the average annual real return on US equities at about 7% and on short-term government debt at about 1% over 1889–1978, leaving a premium of roughly 6 percentage points. They then asked what premium their representative-agent consumption-CAPM with CRRA preferences could generate, feeding in the observed smoothness and weak equity-correlation of US consumption growth. Restricting the risk-aversion coefficient to the behaviorally defensible range (they capped it at 10), the model could produce a premium of at most about 0.35 percentage points — roughly a factor of seventeen too small. To reproduce the actual 6 points the model needs risk aversion around 30–40, far outside any independently plausible value. That order-of-magnitude shortfall is the puzzle.
Mapped back: The consumption-CAPM with CRRA preferences is the calibrated model; the ~6-point 1889–1978 premium is the observed quantity; the risk-aversion cap of 10 is the independently constrained parameter. The model's maximum ~0.35-point output against the 6-point target is the order-of-magnitude gap — too large to be a tuning failure, which is why it indicts a layer of the enumerable assumption stack rather than merely a mis-set coefficient.
Applied / In Practice¶
Robert Barro's 2006 "Rare Disasters and Asset Markets in the Twentieth Century" is a concrete field deployment of the tail-risk branch. Rather than treating consumption as smooth, Barro assembled international macroeconomic data on twentieth-century catastrophes — wars, depressions, and the like — across roughly three dozen countries, identifying many episodes in which real per-capita output or consumption fell sharply (on the order of 15% or more, some far deeper). Calibrating a disaster probability of around 1.7% per year with that empirical distribution of contraction sizes, he showed a standard expected-utility model with moderate risk aversion could rationalize an equity premium near the observed magnitude: equities are shunned because they crater precisely in disasters, and the rarely-realized-but-real tail is missing from the placid postwar US sample.
Mapped back: Barro's model is the calibrated model re-specified along one layer of the enumerable assumption stack — the claim that the postwar consumption sample misses the true tail, which is the four-way branch of resolutions' rare-disaster arm. Using disaster-economy data to close the gap is exactly the differential fingerprint the puzzle predicts for this branch: economies that actually suffered catastrophes should show premia the augmented model can match, so the fix is confirmed where its distinct empirical signature should appear.
Structural Tensions¶
T1: Which leg is anomalous — high equity return versus low bond return (the same gap read from two ends). The intuitive statement of the puzzle points at equities: why do stocks pay six points more than they "should"? Mehra and Prescott's sharpened reading relocates the anomaly to the other leg — the question becomes why the risk-free return is so low that agents who appear only mildly risk-averse won't bid it up, spawning the risk-free-rate puzzle as the same gap seen from the bond side. The tension is that the two framings point resolutions in different directions: a preference fix that raises the model's premium and one that lowers its implied risk-free rate address the same six-point spread but commit to different behavioral claims. An account that fixes only one leg may reopen the other. Diagnostic: Is this resolution closing the gap by rationalizing the equity return, or by rationalizing why agents tolerate so low a bond return?
T2: Persistence versus arbitrage (a large durable spread that is not free money). A six-point spread that has survived ninety years invites the reflex "large, durable mispricing — therefore exploitable." The puzzle's whole difficulty is that the inference fails: the premium persists precisely because it compensates real (if mismeasured) risk, so it is not a riskless profit that trading away would erase. The tension is that the magnitude and durability that make the spread puzzling are the same features that would, in an arbitrage, signal free money — and the discipline is to refuse "large persistent spread, therefore mispricing" and read the spread as unmodeled risk instead. Collapse the distinction and one mistakes a calibration failure of a model for an exploitable inefficiency of a market. Diagnostic: Would trading the spread away erase it (arbitrage), or does it survive because it pays for a risk the model mismeasures (the puzzle)?
T3: The size of the miss as license (recalibration versus structural refutation). A small gap between model and data invites tuning — nudge the risk-aversion coefficient and move on. The equity premium puzzle's force comes from the magnitude: a factor-of-seventeen miss (0.35 points against 6) at any defensible risk aversion is too large to be a mis-set parameter, so it licenses the stronger inference that a load-bearing assumption is false rather than merely mis-calibrated. The tension is that the same number could, in principle, be read either way, and the line between "recalibrate" and "reject the frame" is a judgment about how independently constrained the inputs really are. Push risk aversion to 30–40 and the miss vanishes arithmetically — but only by abandoning the behavioral constraint that made the miss diagnostic in the first place. Diagnostic: Is the gap small enough to absorb by re-tuning a constrained parameter, or large enough that closing it that way violates the constraint that made it a puzzle?
T4: Assumption-localization versus entanglement (clean one-layer bets on a coupled stack). The puzzle's discipline is to read every candidate fix as a wager on exactly one named layer of a short stack — preferences, the consumption tail, the marginal investor, or the data series — so rivals compare on common ground. The convenience of that framing is also its risk: the layers are not fully independent. Epstein-Zin preferences and rare disasters can each shoulder part of the same gap; limited participation changes whose consumption enters the covariance, which interacts with the preference specification. The tension is that "which single assumption does this abandon?" imposes a clean four-way partition on a model whose assumptions co-determine the implied premium, so a resolution that genuinely touches two layers gets forced into one branch and mis-compared. Diagnostic: Does this fix isolate its bet to one layer of the stack, or does it silently re-specify a second layer that another branch also claims?
T5: Differential fingerprints versus four decades without consensus (testable branches, undecided contest). The puzzle earns its keep by telling each branch where to leave an empirical signature — disaster-economy premia for tail risk, disaggregated stockholder consumption for limited participation, lower interrupted-foreign-market premia for survivorship bias. That should make the contest decidable. Yet after forty years none of the resolutions commands consensus. The tension is that the differential-prediction machinery promises adjudication the field has not delivered: each fingerprint is confirmable in isolation — Barro's disaster calibration matches, participation studies find volatile stockholder consumption — without any single branch closing the gap so decisively that the others fall. The puzzle sharpens what each story commits to and still leaves the choice among confirmed stories open. Diagnostic: Does this branch's fingerprint merely appear where predicted, or does its appearance actually exclude the rival branches that also fit the six-point premium?
T6: Depth on one number versus breadth across the subfield (a single anomaly, not a recurring mechanism). The apparatus is rich — one calibration target, an enumerable stack, a four-way branch, differential predictions — but it is deployed against exactly one gap. Unlike a mechanism that reappears as fresh instances across settings, the equity premium puzzle does not recur; there is no second asset-pricing context where the same consumption-CAPM-versus-risk-asset gap shows up as a distinct object. Its reach within the domain is depth on one number, sitting in a small cluster of consumption-CAPM anomalies (risk-free-rate, volatility, value-premium puzzles) that indict the same over-narrow utility frame. The tension is between the generality of the reasoning the puzzle models and the singularity of the object it is about: powerful method, one specimen. Diagnostic: Are you invoking the puzzle as a reusable diagnostic pattern, or as the one specific asset-pricing anomaly it names?
T7: Autonomy versus reduction (its own named puzzle or an instance of the calibration-anomaly parent). The equity premium puzzle is a canonically named, forty-year research object with proprietary furniture — the consumption covariance, CRRA preferences, the 6-point premium, the rare-disaster and limited-participation branches. Yet its portable content sits one level up: it instantiates the general pattern a calibrated model diverges from a measured quantity by more than tuning can absorb, indicting a load-bearing assumption, a mismeasured input, or the wrong frame, together with backward inference from the size of the miss and localization to one layer of an explicit stack. That parent travels across the empirical sciences as co-instances; the EPP's own machinery is asset-pricing furniture that does not and should not travel. Any "X-premium puzzle" elsewhere borrows the problem-statement form by analogy, carrying the parent under a home-domain name. Diagnostic: Resolve toward the calibration-anomaly / model-data-gap parent when asking what reasoning travels beyond asset pricing; toward the named puzzle when diagnosing the specific six-point consumption-CAPM miss in situ.
Structural–Framed Character¶
The equity premium puzzle sits at the framed-leaning position on the structural–framed spectrum — well short of the framed pole occupied by a practice-verdict like ad hominem, but pulled decisively off structure by its dependence on a specific modeling tradition. The five criteria pull unevenly, and the split is what fixes it as framed-leaning rather than mixed.
On evaluative weight it is only mildly charged, and that mildness is the one thing keeping it off the framed pole. "Puzzle" flags an anomaly — something is wrong — but the entry is emphatic that the charge falls on the model, not on the world or on any agent's reasoning: it is "not a verdict that markets are irrational or inefficient" but a calibration failure of one framework. There is no normative conviction of a person or practice the way "ad hominem" convicts a move; the anomaly is an epistemic finding about a model's fit, closer to neutral diagnostic than to blame. That leg leans structural. Every other leg leans framed, and hard. On human-practice-bound the concept is constituted by the practice of consumption-based asset pricing and dissolves without it: strip away the representative-agent consumption-CAPM, the CRRA calibration, and the behaviorally-constrained risk-aversion range, and the bare fact that equities have out-returned bonds is just a return series, with no "puzzle" to it — the gap exists only relative to what a particular model predicts, so remove the modeling practice and there is nothing for the anomaly to grip. On institutional origin it is a dated artifact of a specific theoretical tradition (Mehra–Prescott 1985), inseparable from the consumption-covariance machinery, the four-way branch of resolutions, and the risk-free-rate-puzzle apparatus — all distinctions drawn inside asset-pricing theory, not substrate-neutral form. On vocab_travels it scores low: consumption covariance, CRRA preferences, the marginal investor, rare-disaster tails, and survivorship bias are all pinned to the finance substrate and lose their referents off it. On import_vs_recognize the transfer beyond asset pricing is import-by-analogy, not mechanism-recognition — a "compensation premium puzzle" in labor borrows the problem-statement form while supplying no consumption-CAPM, no calibrated risk parameter, and no return series, so the word does evocative, not analytic, work.
The one structural-looking feature is the calibration-anomaly / model–data-gap skeleton: a calibrated model diverges from an independently measured quantity by more than tuning can absorb, indicting one named layer of an enumerable assumption stack — with backward inference from the size of the miss and localization of the failure to a single layer. That skeleton is genuinely portable and recurs across the empirical sciences as co-instances, which is what tempts a structural reading. But it does not lift the named puzzle off the framed-leaning region, because that skeleton is exactly what the equity premium puzzle instantiates from its parent, not what makes "the equity premium puzzle" itself travel: the cross-domain reach belongs to the general calibration-anomaly pattern, while the puzzle's distinctive content — the consumption covariance, the 6-point premium, the specific four-way branch — is precisely the part that stays home. Its character: a lightly-charged but deeply practice-constituted asset-pricing anomaly whose every distinctive component is modeling-tradition furniture, structural only in the calibration-anomaly skeleton it instantiates from its parent.
Structural Core vs. Domain Accent¶
This section decides why the equity premium puzzle is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity in the same breath — so it is worth being exact about what could lift and what stays home.
What is skeletal (could lift toward a cross-domain prime). Strip the finance and a thin reasoning structure survives: a model whose inputs are independently constrained predicts a quantity that is independently measured, the two diverge by a factor too large for tuning to absorb, and the size of the miss indicts one named layer of an enumerable stack of premises rather than a mis-set knob. The portable pieces are abstract — a calibrated predictor, a measured target, a gap whose magnitude does epistemic work (a small gap invites recalibration, an order-of-magnitude gap licenses the stronger inference that a load-bearing assumption is false), a short enumerable assumption stack, and the discipline of reading every candidate resolution as a wager on exactly one layer with its own differential empirical fingerprint. That skeleton is genuinely substrate-portable — it is the calibration-anomaly / model–data-gap pattern the entry names as its parent, and it recurs across the empirical sciences wherever an independently-constrained model misses an independently-measured target by orders of magnitude. But it is the core the puzzle shares with those co-instances, not what makes the equity premium puzzle the specific object it is.
What is domain-bound. Almost every distinctive component is asset-pricing furniture that does not survive extraction. The failing model is a specific one — the representative-agent consumption-CAPM with CRRA preferences, which prices the premium as the price of the covariance of equity returns with aggregate consumption growth. The measured target is a specific number — the ~6-percentage-point excess return of US equities over short-term government bonds across 1889–1978. The constrained parameter is a specific coefficient — risk aversion pinned to 1–10 by behavioral evidence, against the 30–40 the model needs. The four-way branch of resolutions (preference-side habit/Epstein-Zin/loss aversion, rare-disaster tails, limited participation, survivorship bias), the risk-free-rate puzzle as its mirror leg, and the differential fingerprints each branch is tested against are all distinctions drawn inside asset-pricing theory. The decisive test: remove the consumption-CAPM and its behaviorally-calibrated risk parameter and the residue is just a return series — equities have out-returned bonds — with no anomaly in it at all, because the "puzzle" exists only relative to what that particular model predicts. The gap is constituted by the very modeling practice the prime bar asks it to shed.
Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. The equity premium puzzle's transfer is bimodal, and both modes keep it below the bar. Within asset pricing it does not even transfer as a recurring mechanism: it is one anomaly, depth on a single number, sitting in a small cluster of consumption-CAPM anomalies (risk-free-rate, volatility, value-premium) that indict the same over-narrow utility frame — there is no second setting where the same gap reappears as a fresh instance. Beyond asset pricing it travels only by analogy: a "compensation premium puzzle" in labor, or any other "X-premium puzzle," borrows the problem-statement form while supplying no consumption-CAPM, no calibrated risk parameter, and no return series, so the name does evocative rather than analytic work. When the bare structural lesson — a calibrated model diverging from a measured quantity beyond tuning, indicting one layer of an explicit stack — is genuinely needed cross-domain, it is already carried, in more general form, by the calibration-anomaly / model–data-gap parent the entry instantiates. The cross-domain reach belongs to that parent; the named puzzle carries consumption-covariance, CRRA, the 6-point premium, and the rare-disaster/limited-participation/survivorship branches, which are precisely the baggage that should stay home.
Relationships to Other Abstractions¶
Current abstraction Equity premium puzzle Domain-specific
Parents (1) — more general patterns this builds on
-
Equity premium puzzle is a decomposition of Calibration Anomaly Prime
The Equity Premium Puzzle is the asset-pricing form of a calibration anomaly in which an independently constrained model misses an independently measured target beyond plausible retuning.Removing the consumption-CAPM, CRRA, and historical return-series vocabulary leaves a quantitative theory-observation gap that survives noise and forces diagnosis of a short assumption stack. Calibration Anomaly carries that portable reasoning pattern; the six-point equity spread and its resolution branches are the domain frame.
Children (1) — more specific cases that build on this
-
Risk-Free Rate Puzzle Domain-specific presupposes Equity premium puzzle
The Risk-Free Rate Puzzle arises when raising CRRA risk aversion to repair the Equity Premium Puzzle drives the same model's risk-free-rate prediction implausibly high.Weil's companion anomaly is defined by what the natural repair of the Mehra–Prescott miss does to a second empirical target. Without the first puzzle's demand for a much larger equity premium, there is no reason to turn the shared risk-aversion parameter high enough to generate this distinctive rate-level failure. The domain-to-domain edge preserves that sequence and avoids flattening both puzzles directly under the same generic calibration node.
Hierarchy path (1) — routes to 1 parentless root
- Equity premium puzzle → Calibration Anomaly
Not to Be Confused With¶
-
The equity risk premium (the quantity). The equity risk premium is simply the measured excess return of equities over safe debt — the ~6-point number itself, an empirical fact about a return series. The puzzle is not that number but the gap between it and what a consumption-CAPM with defensible risk aversion can rationalize; the premium exists and is uncontroversial, while the puzzle is the anomaly it poses for a particular model. Tell: is the object a return spread you could measure with no model in hand (the premium), or a discrepancy that only exists relative to what a calibrated model predicts (the puzzle)?
-
The risk-free-rate puzzle. Its mirror sibling — the same six-point gap read from the bond leg: why is the safe return so low that agents who appear only mildly risk-averse won't bid it up? It is not a separate anomaly but the identical spread seen from the other end, so a preference fix that raises the model's equity premium and one that lowers its implied risk-free rate address the same gap while committing to different behavioral claims. Tell: is the resolution rationalizing why equities pay so much (equity-premium leg) or why bonds pay so little (risk-free-rate leg)?
-
Sibling consumption-CAPM anomalies (excess-volatility, value-premium puzzles). Other calibration failures of the same over-narrow utility framework — Shiller's finding that prices swing more than dividend fundamentals justify, or the excess return of value over growth stocks. They cluster with the equity premium puzzle and indict the same representative-agent apparatus, but each is a distinct gap against a distinct moment of the data. Tell: which measured quantity is the model missing — the equity-vs-bond mean spread (this entry), price volatility, or the value-growth cross-section?
-
The consumption CAPM itself. The representative-agent, CRRA, consumption-covariance model is the calibrated model whose failure constitutes the puzzle — it is the machinery under indictment, not the anomaly. Confusing the two treats the framework as if it were the finding. Tell: are you naming the pricing model that predicts the premium (CCAPM), or the order-of-magnitude miss between its prediction and the data (the puzzle)?
-
Behavioral-mispricing / limits-to-arbitrage anomalies. A different family of asset-pricing puzzles — bubbles, momentum, closed-end-fund discounts — that indict the market as irrational or inefficient and rest on frictions that stop arbitrageurs from correcting it. The equity premium puzzle indicts the model, not the market: the premium persists as compensation for real (if mismeasured) risk, not as an exploitable mispricing. Tell: does the anomaly say investors are behaving irrationally and a trade would erase it (mispricing), or that a specific pricing model is misspecified while the spread rationally survives (this entry)?
-
The calibration-anomaly / model–data-gap pattern (parent). The broader structure the puzzle instantiates — a calibrated model diverging from an independently measured quantity by more than tuning can absorb, indicting one named layer of an enumerable assumption stack — which recurs across the empirical sciences as co-instances. The equity premium puzzle is the asset-pricing specimen keyed to consumption covariance and a return series; the parent is the reusable reasoning pattern. Tell: are you invoking a portable "calibrated model misses measured target → localize the failing assumption" method (the parent, treated more fully elsewhere), or the one specific consumption-CAPM miss (this entry)?
-
Analogical "X-premium puzzle" coinages. Borrowings like a "compensation premium puzzle" in labor economics reuse the problem-statement form — a stubborn premium a standard model can't rationalize — but supply no consumption-CAPM, no calibrated risk-aversion parameter, and no equity return series. The name travels by analogy; the object does not. Tell: does the coinage carry the actual consumption-covariance machinery (the same puzzle), or only the rhetorical shape of "a premium our model can't explain" (analogy)?
Neighborhood in Abstraction Space¶
Equity premium puzzle sits in a crowded region of the domain-specific corpus (35th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Macroeconomic Equilibria & Consumer Demand (19 abstractions)
Nearest neighbors
- Risk-Free Rate Puzzle — 0.92
- St. Petersburg Paradox — 0.84
- Modigliani–Miller theorem — 0.84
- Income Elasticity of Demand — 0.84
- Giffen Good — 0.83
Computed from structural-signature embeddings · 2026-07-12