Skip to content

Beauty Contest Game

Have players simultaneously pick a number to land closest to p times the group average, so that iterated best-response contracts the target toward zero — turning the infinite tower of 'what others expect others to expect' into a single measurable reasoning-depth scalar read off the choices.

Core Idea

The beauty contest game — Keynes's 1936 metaphor, operationalized experimentally by Nagel (1995) — is a coordination game in which each of n players simultaneously picks a number in some bounded range, and the payoff rewards the player whose number is closest to p times the average of all choices, where p < 1. The game is designed so that iterating the best-response logic from any prior belief contracts the target: if others are expected to pick the midpoint, one should pick p times that, and if others are expected to do the same, one should pick p² times it, and so on — reaching the unique Nash equilibrium at zero by iterated dominance. The game is therefore a clean probe of iterated higher-order belief: how many rounds of "I think that you think that I think..." a real population actually executes before stopping.

The mechanism Keynes identified and Nagel formalized is level-of-reasoning depth as an observable quantity. Level-0 behavior is picking randomly or uniformly; level-1 is best-responding to level-0 (pick p × midpoint); level-2 is best-responding to level-1; and so on. Each level generates a distinct prediction for what players choose, and the distribution of choices in a population reveals the distribution of reasoning levels. Laboratory results are consistent: in first-play experiments with p = ⅔, the modal response clusters near 33 (level-1) and 22 (level-2); very few subjects choose 0; the Nash prediction requires not just individual rationality but common knowledge of rationality at all orders, a condition empirically unmet. The name derives from Keynes's newspaper contest analogy — a reader who wins by guessing which face other readers will vote for must guess not who is prettiest but who others will think others will think prettiest — and the extension to asset pricing: stock prices under speculative trading reflect not fundamental value but what each trader expects others to expect about future prices, a regression that terminates at finite depth rather than converging to fundamentals.

Structural Signature

Sig role-phrases:

  • the simultaneous players — n agents each picking a number in a bounded range at once
  • the p-times-average winning rule — the payoff rewards proximity to p times the average of all choices, with p < 1
  • the contraction toward zero — iterating best response from any prior belief multiplies the target by p each round, driving the unique Nash equilibrium to 0
  • the level-of-reasoning depth — the observable scalar k: level-0 picks randomly, level-1 best-responds to level-0, level-2 to level-1, and so on
  • the depth-to-choice mapping — reasoning depth k maps to a distinct predicted choice ≈50·pᵏ, so the choice distribution reveals the depth distribution
  • the finite-depth stopping — empirical play clusters at level-1 (~33) and level-2 (~22), with almost no one choosing 0
  • the common-knowledge gap — the equilibrium at 0 requires common knowledge of rationality at every order, which the finite stopping shows is empirically unmet
  • the level-shifting conditions — depth rises with stakes, experience, and expertise, and repeated play converges toward equilibrium
  • the best-respond-one-level-deeper strategy — optimal play sets one's own level just above the counterparty's measured level, not at the infinite-depth equilibrium

What It Is Not

  • Not evidence of irrationality. Play stopping at level-1 (~33) or level-2 (~22) rather than the Nash equilibrium at 0 is not a failure of individual rationality. It is a measurement of how far up the tower of mutual confidence — "everyone is rational, and everyone is confident everyone is rational, without end" — a real population climbs. The gap from equilibrium localizes precisely to the common-knowledge-of-rationality assumption, not to rationality as such.
  • Not literally about beauty or aesthetics. Keynes's newspaper-contest analogy is a metaphor: the winner guesses not who is prettiest, nor whom others find prettiest, but whom others think others will think prettiest. The game is a probe of iterated higher-order belief; aesthetics is the illustrative dressing, not the subject.
  • Not a Bayesian Nash / incomplete-information game. It is a complete-information game with no private types. The quantity of interest is bounded-rationality depth — how many rounds of best-response a population executes — not equilibrium computation under a prior over types.
  • Not a classical coordination game with multiple equilibria. Coordination games have several equilibria and ask which is selected. The beauty contest has a unique Nash equilibrium (0, by iterated dominance) and asks how close to it human play actually gets. The puzzle is depth of reasoning, not equilibrium selection.
  • Not a game where optimal play is the Nash equilibrium. Choosing 0 loses against a population reasoning at finite depth. The optimal action is to reason one step deeper than the counterparty's measured level — p times the expected average given that level — not to assume the infinite-depth equilibrium.
  • Not a portable mechanism, and its finance use is a metaphor. The depth estimate is an instrument valid only among strategically reasoning agents; the transferable ideas (bounded iteration depth, higher-order-belief dependence, unwarranted common knowledge) belong to common_knowledge, bounded_rationality, and expectations. Even Keynes's speculative-pricing extension — prices tracking expectations-about-expectations — travels as the "Keynesian beauty contest" metaphor, importing the level-k reading by analogy, not deploying the experimental paradigm.

Scope of Application

The beauty contest game lives within a single home discipline — behavioural game theory — restaged across its strategic-interaction settings; its reach is bounded to populations of agents running iterated best response on a complementarity payoff (the depth estimate is an instrument valid only where such agents exist), and the general ideas it makes vivid (bounded iteration depth, higher-order-belief dependence, unwarranted common knowledge) are carried by the parents common_knowledge, bounded_rationality, coordination_problem_and_equilibrium_selection, and expectations.

  • Experimental economics — the home turf: the canonical workhorse for measuring cognitive-hierarchy / level-k reasoning across cultures, expertise levels, and incentive structures.
  • Macroeconomic expectations — the beauty-contest payoff (utility decreasing in distance from the average action) is the analytic core of Morris-Shin and Angeletos-style coordination-under-public-information models.
  • Behavioural finance — the "Keynesian beauty contest" theory of speculative pricing, bubbles, and momentum, where traders price against expected future beliefs of others rather than fundamentals (a metaphorical extension, importing the level-k reading by analogy).
  • Bounded-rationality modelling of strategic games — the level-k apparatus born here exported to auctions, signalling games, market-entry games, and policy-experimentation settings.

Clarity

The beauty contest game's clarifying force is that it turns "how many rounds of strategic recursion do people actually run?" — a question that had been gestured at philosophically but never measured — into a single observable number. Because each reasoning level k maps to a distinct predicted choice (≈50·pᵏ), the distribution of choices in a population becomes a direct read-out of the distribution of reasoning depths: a cluster at 33 is level-1, a cluster at 22 is level-2, and the near-absence of choices at 0 says that almost no one carries the iteration to its limit. What was an untestable claim about the depth of "I think that you think that I think…" becomes a quantity an experiment can estimate, compare across populations, and watch move with stakes and experience.

The sharper distinction the game makes legible is between individual rationality and common knowledge of rationality at every order. The Nash equilibrium at 0 is reachable by iterated dominance, but only for players who are not merely rational themselves but confident that everyone is rational, and that everyone is confident that everyone is rational, without end. Empirical play stopping at level-1 or level-2 is therefore not evidence of irrationality; it is a measurement of exactly how far up that tower of mutual confidence a real population climbs before it stops — and the gap between the equilibrium prediction and the observed mean localizes the failure precisely to the higher-order belief assumption rather than to rationality as such. This lets the analyst pose the operative question — at what level is my counterparty reasoning, and at what level should I therefore reason? — rather than assuming the infinite-depth equilibrium, which is what makes the game the workhorse probe for cognitive-hierarchy models and the analytic core of the Keynesian view that speculative prices track expectations-about-expectations rather than fundamentals.

Manages Complexity

The complexity the beauty contest game tames is the unbounded recursion of higher-order belief. "What others expect others to expect" extends in principle without limit — belief about belief about belief, an infinite tower — and that tower is the thing that actually drives behavior in speculative markets and coordination settings, yet as stated it is intractable: there is no obvious way to say how far up it any given population, trader, or experimental subject actually climbs. The game collapses the whole tower onto a single scalar, the level of reasoning k, by exploiting a structural fact about its payoff: iterating best response from any starting belief multiplies the target by p each round, so reasoning depth k maps to one predicted choice, ≈50·pᵏ. An infinite-dimensional object — the full hierarchy of beliefs about beliefs — is thereby parameterized by a single number, and the analyst tracks just that number instead of the recursion.

Because each depth produces a distinct, separated prediction, the read-off is immediate and the branch structure is a discrete ladder. A choice near 33 is level-1, a choice near 22 is level-2, the next cluster level-3, and the near-absence of choices at 0 says almost no one runs the recursion to its limit; the population's whole distribution of strategic sophistication is read straight off the distribution of choices, one number per player. From that same scalar the analyst reads several things at once: the typical depth a given population reaches; how the depth shifts with the conditions that move it — rising with stakes, with experience under repeated play, with expertise — so that heterogeneous populations are compared on one axis rather than re-modeled each time; and the precise location of the equilibrium's failure, since the gap between the Nash prediction at 0 and the observed mean localizes entirely to the higher-order-belief assumption (common knowledge of rationality at every order) rather than to individual rationality. So a problem that began as an infinite epistemic regress with no measurable handle becomes: collect one number per player, place it on the pᵏ ladder, and read off reasoning depth, its responsiveness to stakes and experience, and the exact order at which mutual confidence gives out — the recursion compressed to a single estimated quantity with a clean discrete branch.

Abstract Reasoning

The beauty contest game licenses inferences that turn an infinite epistemic regress into a measurable scalar and a strategy keyed to it.

Diagnostic — read reasoning depth off the choices. The signature move is to infer a player's (or a population's) level of reasoning directly from the number chosen, using the structural fact that depth k maps to the predicted choice ≈50·pᵏ. A choice near 33 is read as level-1, a choice near 22 as level-2, the next cluster as level-3, and the near-absence of choices at 0 as evidence that almost no one runs the recursion to its limit. So the analyst reasons from an observed action distribution to the distribution of strategic sophistication in the population — one number per player, placed on the pᵏ ladder — converting an untestable claim about "I think that you think…" into an estimated quantity.

Strategic best-response — set your own level one above your counterparty's. The actionable move follows: having inferred the level at which the counterparty is reasoning, the analyst chooses to reason one step deeper, best-responding to the estimated level rather than to the infinite-depth equilibrium. The inference is explicitly not to play the Nash prediction of 0 — that loses against a population reasoning at finite depth — but to play p times the expected average given the counterparty's measured level. The operative question is therefore "at what level is my counterparty reasoning, and at what level should I therefore reason?", and the optimal action is read off the counterparty's depth, not assumed from equilibrium.

Boundary-drawing — localize the equilibrium's failure to higher-order belief. A central interpretive move is to attribute the gap between the Nash prediction (0) and observed play correctly. The equilibrium at 0 is reachable by iterated dominance but requires not just individual rationality but common knowledge of rationality at every order — that everyone is rational, and everyone is confident everyone is rational, without end. The analyst infers that play stopping at level-1 or level-2 is therefore not evidence of irrationality but a measurement of exactly how far up that tower of mutual confidence the population climbs before stopping, so the failure is localized precisely to the higher-order-belief assumption rather than to rationality as such. This bounds when the infinite-depth equilibrium can be assumed at all: only where common knowledge of rationality is empirically warranted, which the game shows it generally is not.

Comparative-statics prediction — what moves the level. Because reasoning depth is a single quantity, the analyst predicts how it shifts with conditions: it rises with stakes, with experience under repeated play, and with expertise (professional traders play deeper than undergraduates), and repeated play with the same population converges toward 0. So heterogeneous populations are compared on one axis, and the analyst forecasts that raising incentives or letting a population learn will pull observed choices down the pᵏ ladder toward equilibrium — without re-modeling each population.

Reframing — speculative prices as expectations about expectations. A further move extends the structure to asset pricing: the analyst reasons that under speculative trading a price reflects not fundamental value but what each trader expects others to expect about future prices, a regression that terminates at finite depth rather than converging to fundamentals. The inference is that bubbles and momentum can be read as the market settling at a shallow reasoning level about others' beliefs, so the analyst looks for the operative depth of belief-about-belief rather than assuming prices track fundamentals.

Knowledge Transfer

Within behavioural game theory the beauty contest game transfers as mechanism, carried by the level-k framework it operationalised. The diagnostic (read a player's or population's reasoning depth off the chosen number using the depth-k → ≈50·pᵏ mapping), the strategic prescription (set your own level one step above your counterparty's measured level rather than playing the infinite-depth equilibrium), the boundary-drawing move (localise the gap between the Nash prediction at 0 and observed play to the common-knowledge-of-rationality assumption, not to individual rationality), and the comparative statics (depth rises with stakes, experience, and expertise, and repeated play converges toward equilibrium) all carry intact wherever strategic complementarities make higher-order beliefs about others' actions determine optimal play. So the level-k apparatus born here exports across the home domain to auctions, signalling games, market-entry games, and policy-experimentation settings as a bounded-rationality model, and the beauty-contest payoff structure (utility decreasing in distance from the average action) is the analytic core of the Morris-Shin and Angeletos-style macroeconomic models of coordination under public information. These are all strategic-interaction settings — one substrate restaged — so this is reach within a domain, not transfer across substrates.

Beyond strategic interaction the honest characterisation has two threads, and neither is "the beauty contest game as a portable mechanism." First, the game is in part a measurement probe: its real output is the estimate of reasoning depth, and that estimate is meaningful only where its precondition holds — a population of agents running iterated best response on a complementarity payoff. Where there are no strategically reasoning agents, there is no reasoning depth to read, so the probe does not extend; the boundary to mark is instrument-reach, not metaphor. Second, the underlying structural ideas the game makes vivid — that iterated reasoning has bounded depth, that strategic complementarities create higher-order-belief dependencies, that common-knowledge-of-rationality is empirically unwarranted — are genuinely general, but they are carried by the parent primes, not by the beauty contest itself: common_knowledge (the epistemic concept the game probes finite-depth deviations from), bounded_rationality (the depth limit it measures), coordination_problem_and_equilibrium_selection (the strategic-complementarity frame), and expectations (beliefs-about-beliefs as the driver). When the cross-domain lesson is needed — "do not assume agents run the recursion to its limit; ask how deep they actually go" — it should be carried by those parents, of which the beauty contest is the canonical experiment, not an additional structural pattern. Notably, even Keynes's own most famous extension — that speculative asset prices reflect what each trader expects others to expect rather than fundamentals — is offered and travels as a metaphor (the "Keynesian beauty contest" theory of bubbles and momentum), illuminating by resemblance and supplying a vivid frame, but importing the level-k reading by analogy rather than deploying the experimental paradigm's machinery. So the honest move is to attribute the portable content to the higher-order-belief and bounded-rationality primes, to treat the depth estimate as an instrument valid only among reasoning agents, and to flag the finance application as the metaphor it has always been (see Structural Core vs. Domain Accent).

Examples

Canonical

Rosemarie Nagel's 1995 experiment is the defining operationalisation. Subjects each chose a number from 0 to 100, and whoever landed closest to two-thirds (p = ⅔) of the group average won a fixed prize. Work the iterated best response: if others choose uniformly, the average is 50, so a level-1 reasoner should pick ⅔ × 50 ≈ 33; a level-2 reasoner, expecting others to pick 33, should pick ⅔ × 33 ≈ 22; carried to its limit the target contracts to the unique Nash equilibrium of 0. Nagel's first-play data showed exactly the shallow-depth signature: pronounced clusters of choices near 33 and near 22, and almost no one at 0. The choice distribution thereby measured the population's distribution of reasoning depth directly.

Mapped back: The number-picking with the two-thirds rule is the simultaneous players under the p-times-average winning rule, whose iterated best response is the contraction toward zero. The spikes at 33 and 22 are the depth-to-choice mapping (≈50·pᵏ) made visible as the finite-depth stopping, and the empty equilibrium at 0 is the common-knowledge gap.

Applied / In Practice

In 1997 Richard Thaler ran the game as a mass field experiment in the Financial Times: readers were invited to guess an integer from 0 to 100, with the entry closest to two-thirds of all entries' average winning a prize, drawing on the order of a thousand-plus submissions from a sophisticated readership. The average of the entries came to roughly 18.9, so the winning guess — two-thirds of that — was 13. The result placed the crowd's effective reasoning at about level 2 to 3 rather than at the equilibrium 0: many entrants clearly ran one or two rounds of "others will pick 33, so I pick 22, so I pick lower," but the recursion stopped well short of its limit, and a scatter of naive high guesses and a spike at 0 (over-clever entrants) framed the sophisticated middle.

Mapped back: The FT readers are the simultaneous players under the p-times-average winning rule; the winning 13 sits low on the ≈50·pᵏ ladder of the depth-to-choice mapping, indicating the finite-depth stopping near level 2–3. That a real, motivated crowd still did not reach 0 is the common-knowledge gap, and the readership's sophistication illustrates the level-shifting conditions (expertise pulls choices down the ladder).

Structural Tensions

T1: Unique equilibrium versus the equilibrium being the losing move (rationality that must not be played). The game has a clean, unique Nash equilibrium at 0, reachable by iterated dominance — the tidy answer game theory prizes. Yet choosing 0 reliably loses against any real population reasoning at finite depth, so the equilibrium action is precisely the wrong action. This is the game's defining paradox: correct equilibrium analysis and optimal play point in opposite directions, and the skilled player must best-respond to the counterparty's measured depth rather than to the theory's prescription. The tension is that the game simultaneously honours the equilibrium concept (it exists, it is derivable) and refutes its behavioural authority (playing it is a mistake), so it is both a showcase for iterated dominance and a demonstration that iterated dominance is the wrong guide to action against boundedly-rational opponents. Diagnostic: Is the recommended action the infinite-depth equilibrium (correct as analysis, losing as play) or one step above the counterparty's actual reasoning depth (the winning move)?

T2: Depth as a clean scalar versus the level-0 anchor it hides (an identification problem inside the ladder). The pᵏ ladder collapses the infinite belief hierarchy to one number by mapping depth k to ≈50·pᵏ. But that mapping is anchored on an assumed level-0 (uniform random, or midpoint), and the same observed choice is consistent with different combinations of reasoning depth and beliefs about what level-0 others are doing — a level-2 reasoner with one anchor can land where a level-1 reasoner with another does. The tension is that the scalar's cleanliness depends on fixing the level-0 specification, which is itself a modelling choice the data underdetermine, so "this player is level-2" is partly an artifact of the anchor the analyst assumed. The compression that makes reasoning depth measurable also imports an identification assumption that the single number conceals. Diagnostic: Is the inferred reasoning depth robust to the assumed level-0 anchor, or does the classification of a given choice flip under a different but equally defensible specification of naive behaviour?

T3: Laboratory probe versus ecological validity (a measured depth that may not travel to the market). The game's real output is an estimate of reasoning depth, and within the number-guessing task that estimate is crisp. But the depth measured on an artificial p-times-average task need not be the depth agents run in the messy strategic settings the game is invoked to illuminate — a trader pricing against others' expectations faces vastly richer information, feedback, and stakes than an FT reader guessing an integer. The tension is that the instrument's precision is bought by an artificial task, and the very features that make it a clean probe (bounded range, transparent payoff, single shot) are what make its generalization to real markets uncertain. So the depth number is trustworthy about the game and only conjectural about the phenomena — bubbles, coordination — for which the game is a stand-in. Diagnostic: Is the reasoning depth being used to describe behaviour in the game (valid) or extrapolated to a real strategic setting whose information and incentives differ from the lab task (unwarranted)?

T4: A probe that measures depth versus a probe that erases it (the instrument self-destructs under learning). The comparative statics are a genuine strength: depth rises with stakes and expertise and converges toward 0 under repeated play. But that convergence is also the instrument's undoing — as a population learns, everyone's choices slide down the ladder toward the equilibrium, so the choice distribution loses its discriminating spread and the game stops measuring depth precisely when the population has become sophisticated. The tension is that the same responsiveness which lets the probe track how depth shifts with conditions guarantees that the probe degrades with the very learning it detects: a fresh, naive population yields a rich depth signal, an experienced one collapses toward a single spike at 0 that reveals little. The game is a clean probe only on first contact. Diagnostic: Is the population being measured naive enough for the choice distribution to spread across levels, or has repeated play compressed everyone toward 0, leaving the probe unable to resolve depth?

T5: Autonomy versus reduction (a canonical experiment or the higher-order-belief and bounded-rationality primes it probes). The beauty contest game is a named, canonical paradigm with proprietary apparatus — the p-times-average rule, the ≈50·pᵏ ladder, the 33/22 clusters, Nagel's and Thaler's experiments — but, unusually, it is a measurement instrument rather than a portable mechanism. The depth estimate is meaningful only where its precondition holds: strategically reasoning agents running iterated best response on a complementarity payoff. The structural ideas it makes vivid are genuinely general but are carried by parent primes, not the game: common_knowledge (the epistemic condition it probes finite-depth deviations from), bounded_rationality (the depth limit it measures), coordination_problem_and_equilibrium_selection (the complementarity frame), and expectations (beliefs-about-beliefs as driver). Even Keynes's own speculative-pricing extension travels as the "Keynesian beauty contest" metaphor, importing the level-k reading by analogy. The tension is between a legitimately named experimental paradigm and the recognition that its transferable content belongs to those parents, of which the game is the canonical experiment, not an additional pattern. Diagnostic: Resolve toward common_knowledge + bounded_rationality + expectations when carrying the lesson "do not assume agents run the recursion to its limit"; toward the beauty contest game when the object is an actual population of agents whose reasoning depth is to be measured on a complementarity payoff.

Structural–Framed Character

The beauty contest game sits toward the structural side — best read as mixed-structural, at the domain-bound edge of that band, because it is a neutral formal game and measurement instrument rather than a causal mechanism, yet a specific named paradigm that probes parent primes and reaches beyond reasoning agents only by metaphor. The five criteria lean structural with the pull toward domain-boundedness in transfer. On evaluative weight it reads structural: the game renders no verdict — it explicitly insists that finite-depth play is "not evidence of irrationality" but a measurement of how far a population climbs the tower of mutual confidence. On human-practice-bound it reads mostly structural but agent-bound: the probe requires "a population of agents running iterated best response on a complementarity payoff," so it applies to any strategically reasoning agents (undergraduates, professional traders, an FT readership) rather than to a particular human institution, but it does presuppose cognizing agents with higher-order beliefs and yields nothing without them. On institutional origin it reads structural: it is a named paradigm (Keynes 1936, Nagel 1995), a defined game plus experimental instrument, not an artifact of a survey or agency. On vocab-travels it is mixed: the operative apparatus — the p-times-average rule, the ≈50·pᵏ ladder, level-k depth, the 33/22 clusters — travels intact across auctions, signalling games, and Morris-Shin-style macro coordination models, but the entry stresses these are "one substrate restaged," and the machinery does not extend beyond strategic-complementarity settings. On import-vs-recognize the profile is within-substrate recognition: genuinely the same mechanism across behavioural-game-theory applications, but beyond reasoning agents there is no depth to read, and even Keynes's own speculative-pricing extension "travels as the 'Keynesian beauty contest' metaphor... importing the level-k reading by analogy."

Here the portable structural skeleton is a composition the entry demonstrably needs: the general lesson that iterated reasoning has bounded depth — do not assume agents run the recursion to its limit; ask how deep they actually go — housed in common_knowledge (the epistemic condition it probes finite-depth deviations from), bounded_rationality (the depth limit it measures), coordination_problem_and_equilibrium_selection (the strategic-complementarity frame), and expectations (beliefs-about-beliefs as driver). That lesson is genuinely substrate-spanning, but it is exactly what the beauty contest game instantiates from those umbrella primes, not a force the named game carries on its own — the entry is explicit that "the structural ideas it makes vivid are genuinely general but are carried by the parent primes, not by the beauty contest itself," of which the game is "the canonical experiment, not an additional pattern." So the cross-domain reach belongs to those parents, while the domain-accented apparatus — the p-times-average rule, the pᵏ ladder, the level-k measurement instrument, the depth-shifting comparative statics — stays bound to strategically reasoning populations. Its character: an evaluatively neutral formal game and measurement probe recognised as the same instrument across behavioural game theory, structural in skeleton yet a specific named paradigm that measures its parent primes and reaches non-agents only by metaphor — mixed-structural, at the domain-bound edge, and short of a prime because its portable content is the parents' it probes.

Structural Core vs. Domain Accent

This section settles why the beauty contest game is a domain-specific abstraction and not a prime — a case complicated by its being a measurement instrument rather than a causal mechanism.

What is skeletal (could lift toward a cross-domain prime). Strip away the experimental paradigm and a thin relational lesson survives: iterated reasoning about what others expect has bounded depth — do not assume agents run the recursion to its limit; ask how deep they actually go, because a unique equilibrium reachable by iterated dominance can require an order of mutual confidence a real population never climbs to. The portable pieces are abstract: an infinite tower of belief-about-belief, a contraction that in principle drives it to a limit, and an empirical stopping point short of that limit where common knowledge gives out. That skeleton is genuinely substrate-portable, which is exactly why the entry names it as a composition it demonstrably needs — common_knowledge (the epistemic condition the game probes finite-depth deviations from), bounded_rationality (the depth limit it measures), coordination_problem_and_equilibrium_selection (the strategic-complementarity frame), and expectations (beliefs-about-beliefs as the driver). But this is the core the beauty contest instantiates from those parents, not what makes the named game distinctive.

What is domain-bound. Almost all of the operative content is behavioural-game-theory apparatus, and none of it survives extraction intact: the p-times-average winning rule with p < 1, the contraction toward zero by iterated dominance, the ≈50·pᵏ depth-to-choice ladder that maps reasoning level to a predicted number, the level-k measurement instrument that reads a population's sophistication straight off its choice distribution, the diagnostic 33/22 clusters, the best-respond-one-level-deeper strategy, and the depth-shifting comparative statics (level rising with stakes, experience, and expertise). These are the worked machinery, the instruments, and the empirical cases (Nagel's 1995 experiment, Thaler's 1997 Financial Times contest) the paradigm actually studies, all specific to populations of agents running iterated best response on a complementarity payoff. The decisive test: remove the strategically reasoning agents and there is no reasoning depth to read — the probe does not become a looser probe, it has nothing to bite on and simply does not apply, because its whole output is an estimate that presupposes cognizing agents with higher-order beliefs.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. The beauty contest's transfer is bimodal. Within behavioural game theory the level-k apparatus born here travels intact across auctions, signalling games, market-entry games, and Morris-Shin/Angeletos-style macro coordination models — the same p-average payoff, the same pᵏ ladder, the same comparative statics — because these are one strategic-interaction substrate restaged, not distinct substrates. Beyond strategically reasoning agents it travels only by metaphor: even Keynes's own speculative-pricing extension — prices tracking expectations-about-expectations rather than fundamentals — is offered as the "Keynesian beauty contest" analogy, importing the level-k reading by resemblance rather than deploying the experimental machinery. And when the bare structural lesson is wanted cross-domain — do not assume the recursion runs to its limit; ask how deep belief-about-belief actually goes — it is already carried, in more general form, by the parents the game probes. The cross-domain reach belongs to common_knowledge, bounded_rationality, coordination_problem_and_equilibrium_selection, and expectations; the beauty contest is the canonical experiment that measures those primes in one substrate, and its named apparatus carries game-theoretic baggage that should stay home.

Relationships to Other Abstractions

Local relationship map for Beauty Contest GameParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Beauty Contest GameDOMAINPrime abstraction: Keynesian Beauty Contest — is a decomposition ofKeynesianBeauty ContestPRIMEDomain-specific abstraction: Guess 2/3 of the Average — is a kind ofGuess 2/3 ofthe AverageDOMAIN

Current abstraction Beauty Contest Game Domain-specific

Parents (1) — more general patterns this builds on

  • Beauty Contest Game is a decomposition of Keynesian Beauty Contest Prime

    Removing the bounded-number experiment leaves the Keynesian structure in which the best choice depends on beliefs about others' beliefs and choices.

Children (1) — more specific cases that build on this

  • Guess ⅔ of the Average Domain-specific is a kind of Beauty Contest Game

    Guess ⅔ of the Average is the Beauty Contest Game with p fixed to two thirds, the action interval fixed to 0–100, and its canonical level ladder exposed.

Not to Be Confused With

  • Level-k / cognitive-hierarchy model. The behavioral theory of reasoning depth — level-0 anchors, best-response climbing, a population distribution over levels — that the beauty contest game operationalizes and measures. Part-versus-whole: the model is the account of how depth is structured; the beauty contest is the experimental instrument that reads a depth estimate off choices. The model is applied to many games (entry, auctions, matrix games); the beauty contest is one paradigm where it is measured cleanly. Tell: are you naming the framework that predicts finite-depth reasoning across games (level-k / cognitive hierarchy), or the specific p-times-average guessing task that estimates it (beauty contest game)?

  • Minority game / El Farol Bar problem. A congestion coordination game (Arthur) in which each agent wants to be in the minority — attend the bar only if few others do — so the target is anti-coordination against a capacity threshold, and there is no clean iterated-dominance contraction to a point equilibrium. The beauty contest rewards proximity to p times the average (a complementarity that contracts to 0), not being on the smaller side of a threshold. Tell: is the payoff for landing where few others land (minority/El Farol), or for landing near a shrinking function of where everyone lands (beauty contest)?

  • Centipede game. A sequential game that also probes finite reasoning depth and the empirical failure of common-knowledge-of-rationality, but through backward induction over a series of take-or-pass moves rather than simultaneous number choice; deviation from the game-theoretic prediction there reflects limits on iterated backward reasoning and social preferences, not a level-k depth read off a contraction ladder. Tell: is reasoning depth probed by how far players back-induct in a dynamic take/pass tree (centipede), or by where a population's simultaneous guesses land on the pᵏ ladder (beauty contest)?

  • Iterated deletion of dominated strategies. The solution technique that drives the beauty contest's equilibrium to 0 — repeatedly removing dominated choices until one survives — not the game itself. A reader can mistake the procedure for the paradigm. The game's whole empirical point is that real players stop this iteration early, so the technique yields the equilibrium the players do not reach. Tell: are you naming the general elimination procedure that identifies the unique equilibrium (iterated dominance), or the game whose interest is precisely how far short of that procedure human play stops (beauty contest)?

  • Keynesian beauty contest (speculative-pricing theory). The finance metaphor — that asset prices reflect what traders expect others to expect, so bubbles and momentum track shallow-depth beliefs rather than fundamentals. It shares Keynes's name and the higher-order-belief idea, but it is an interpretive frame imported by analogy, not the controlled experimental paradigm with its measurable pᵏ ladder. Tell: is "beauty contest" being used as a lens on why markets deviate from fundamentals (the pricing metaphor), or as the lab task that estimates reasoning depth from choices (the experimental game)? The entry treats the finance extension as metaphor throughout.

  • The parent primes it probes (common_knowledge, bounded_rationality, expectations). The substrate-neutral concepts the game measures finite-depth deviations from — the epistemic condition, the depth limit, and beliefs-about-beliefs — not confusable peers. Where there are no strategically reasoning agents, there is no depth to read and it is these primes, not "the beauty contest game," that carry the lesson. Tell: strip the guessing task and what remains is the general point that agents do not run the recursion to its limit — carried by common_knowledge / bounded_rationality / expectations, treated fully in the sections above. The game is the canonical experiment that measures them, not the portable structure itself.

Neighborhood in Abstraction Space

Beauty Contest Game sits in a crowded region of the domain-specific corpus (5th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Strategic Interaction & Game Theory (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12