Skip to content

Guess ⅔ of the Average

A single-shot game where each player picks a number closest to two-thirds of the group's mean — its Nash equilibrium of zero is never reached, and because each level of iterated best-response leaves a distinct numerical signature, the modal guess reads off a population's depth of strategic reasoning.

Core Idea

The guess-two-thirds-of-the-average game is a single-shot strategic interaction in which each participant picks a number between 0 and 100 and the winner is the participant whose number is closest to two-thirds of the mean of all submitted numbers. Under common knowledge of rationality and iterated elimination of weakly dominated strategies, the unique Nash equilibrium is zero: no number above 67 can win (two-thirds of the maximum possible average), so anything above 67 is dominated; eliminating those, nothing above 44 can win, and so on until the endpoint of the iterated argument is zero for all players. The game was formalized by Moulin (1986) and given its characteristic experimental shape by Nagel (1995) and the Keynesian-beauty-contest literature.

The structural content is not the Nash prediction but the gap between it and actual play. Real participants perform only a finite number of steps of the iterated-dominance argument. Level-0 players submit around 50 — a focal point in the absence of strategic reasoning. Level-1 players best-respond to a field of level-0 players and submit around 33. Level-2 players best-respond to a field of level-1 players and submit around 22. Experimental populations consistently cluster at levels 1 and 2, well above the equilibrium of zero. The game thereby functions as a measurement instrument for the depth of iterated strategic reasoning in a population: by observing the mean guess, a researcher can infer the modal level of reasoning that subjects perform. In repeated large-scale experiments — newspaper surveys in the Financial Times and Spektrum der Wissenschaft, lab studies with student and professional populations — the instrument has proven reliable, discriminating between populations by education, game-theoretic training, and financial expertise, and providing the primary data against which cognitive-hierarchy models (Camerer, Ho, and Chong) are calibrated.

Structural Signature

Sig role-phrases:

  • the strategic population — participants who each submit a number in a bounded interval (0–100)
  • the recursive winner rule — the prize goes to whoever is closest to a fraction (⅔) of the mean of all submissions, so each player's best choice depends on others' choices
  • the iterated-dominance endpoint — the unique Nash equilibrium of zero reached by chaining "no number above the current ceiling can win," which real play never reaches
  • the level-signature ladder — the engineered reading: each step of iterated best-response maps to one distinct number (level-0 ≈ 50, level-1 ≈ 33, level-2 ≈ 22, level-3 ≈ 14)
  • the modal-guess readout — the instrument's output: the population's mean submission pins its modal depth of strategic reasoning, the recursion rendered as a countable trace
  • the gap-from-Nash datum — the distance between the observed cluster and zero, read as the visible footprint of finite-depth reasoning rather than as failure
  • the parameter levers — the multiplier (⅔ → ½ → p) and the population (training, expertise) shift the per-level signatures and the cluster in known ways, making depth a manipulable dependent variable that calibrates cognitive-hierarchy models
  • the context-laden reading caveat — the limitation: the modal depth is a property of this population under these conditions, not a fixed human constant

What It Is Not

  • Not a test of intelligence or arithmetic. A high guess is not a computation error; the game measures depth of strategic reasoning — how many steps of "best-respond to what others will do" a player performs — not whether they can compute two-thirds of a number. The level signatures (50, 33, 22, 14) track step-count, and clustering at level 2 reflects where the recursion stops, not a failure to do the math.
  • Not a demonstration that subjects are irrational. The gap between the observed cluster and the Nash equilibrium of zero is the visible footprint of finite-depth reasoning, not evidence that players cannot reason. A level-2 guess is a best response to a belief that others reason shallowly; stopping after a few steps because one expects others to stop too is rational, so the gap is data about cognition, not a verdict against it.
  • Not a game where zero is the right number to actually play. Submitting the equilibrium of zero loses against a real population clustered around 22. Zero is the correct target only under common knowledge of rationality, which the game shows fails; the winning play is keyed to the population's expected reasoning depth, so the equilibrium is the benchmark the result departs from, not the recommended action.
  • Not a reading of a fixed human constant. The modal depth is a property of this population under these conditions, not a species parameter: professionals reason deeper, trained students shift downward, and stakes and framing move the number. "People reason at level 2" over-reads a context-laden measurement; the instrument's output is a population reading, not a universal constant.
  • Not a causal mechanism. The ⅔ game is a measurement instrument — a named experimental protocol that reads a hidden cognitive quantity off an observable number — not a force that produces behavior. The level-signature ladder is a readout device, so the concept transfers as an instrument deployed wherever a strategic population exists, not as a mechanism recurring across substrates.

Scope of Application

Because the guess-⅔ game is a named experimental protocol, not a causal mechanism, it applies wherever its precondition can be set up — a population of strategic participants who submit a number under the winner-determination rule — and the fields below are real deployments of the identical instrument, read by the same level-signature ladder. The boundary is instrument-reach versus over-reading: its modal-depth reading is a property of the population under these conditions, not a human constant, and the truncation it reveals travels onward under higher_order_beliefs / bounded_rationality, not as "the ⅔ game."

  • Behavioral game theory / experimental economics — the home laboratory, the canonical instrument for measuring level-k reasoning, run in classrooms, online panels, and lab studies with students and professionals.
  • Cognitive-hierarchy modeling — the workhorse data source whose numerical clusters calibrate Camerer-Ho-Chong-type models in which agents best-respond to a finite-depth distribution of less sophisticated players.
  • Large-scale population surveys — the famous newspaper-readership runs (Financial Times, Spektrum der Wissenschaft), discriminating populations by education, game-theoretic training, and financial expertise.
  • Pedagogical demonstration — classroom and corporate-training use to make higher-order beliefs tangible, and to show distributions shift downward after a game-theory course.
  • Financial-market microstructure (Keynesian beauty contest) — the metaphorical edge: the game is the formal toy of Keynes's beauty contest and the fund-manager-forecasting-others' -forecasts setting, which share the higher-order-belief skeleton but yield no level reading, so the resemblance is analogy rather than the instrument deployed.

Clarity

The game's clarifying contribution is to convert "depth of strategic reasoning" from a slogan into something a behavioural game theorist can read off a single number. Higher-order beliefs — what you think others think others think — are notoriously hard to observe: every recursive layer is internal, and a player's submitted action in most games is consistent with many reasoning depths. The ⅔ game collapses that ambiguity, because each level of iterated best-response maps to a distinct numerical signature (50, 33, 22, 14, …), so the modal guess of a population pins the modal step-count of its reasoning. What had been an untestable hypothesis about cognition becomes a graded measurement, and the gap between the equilibrium of zero and the observed cluster around 22 stops looking like a failure to be rational and starts looking like data — the visible footprint of finite-depth reasoning.

This sharpens the question the field can ask. The naive reading of behaviour against the Nash benchmark is binary and uninformative — players are "irrational" because they miss the equilibrium. The ⅔ game replaces that with a quantitative one: how many steps deep does this population reason, and what moves it? Because the instrument responds to training, expertise, and population, it makes the depth itself the dependent variable, and lets the analyst distinguish two things the bare Nash failure conflates — whether a player cannot perform the iterated-dominance argument at all, versus whether they perform a few steps and rationally stop because they expect others to stop too. That second reading reframes the whole exercise: the right guess is not the equilibrium but a best response to one's belief about others' depth, which is exactly the quantity cognitive-hierarchy models try to pin down, and which the game makes legible by holding the recursion in a form where each layer leaves a countable trace.

Manages Complexity

Higher-order belief is an unbounded recursion — what I expect you to expect me to expect, on without end — and across most strategic settings a single observed action is consistent with many depths of that recursion, so the analyst trying to characterize how deeply a population reasons faces an underdetermined, internal-to-the-head quantity with no clean handle. The ⅔ game compresses that whole tower into a single scalar. Because each step of iterated best-response maps to one distinct number — level-0 at 50, level-1 at 33, level-2 at 22, level-3 at 14, and so on — the recursion is rendered as a labeled sequence, and the modal submission of a population pins its modal step-count directly. An entire family of beliefs-about-beliefs questions collapses to "read the mean guess," and the dependent variable the field cares about (depth of reasoning, and what moves it) becomes one observable number per population rather than an untestable claim about cognition. The same protocol with a different multiplier (½, p, anything) shifts the predicted signatures in a known way, so the analyst varies one parameter and fits cognitive-hierarchy depth across populations from the resulting numerical clusters, rather than building a separate inference for each task. What had been a high-dimensional, recursive, and largely unobservable problem — how far does this population's strategic reasoning go, and how does it respond to training, expertise, or stakes — becomes the tracking of a single count read off a single number, with the gap between the observed cluster and the Nash zero serving as the visible measure of where the recursion truncates.

Abstract Reasoning

The game is, above all, a measurement instrument, and its characteristic moves are those of reading a hidden cognitive quantity off an observable number.

Diagnostic — infer depth of reasoning from the modal guess. The defining inference runs FROM a population's mean (or modal) submission TO the modal number of steps of iterated best-response it performs. Because each level leaves a distinct numerical signature — level-0 at 50, level-1 at 33, level-2 at 22, level-3 at 14 — a cluster around 22 is read as level-2 reasoning, a lower cluster as deeper reasoning. The recursion of higher-order belief, normally internal and unobservable (a submitted action in most games is consistent with many depths), is here pinned to a countable trace, so the analyst infers the unobservable step-count from the observable choice. The gap between the observed cluster and the Nash equilibrium of zero is itself diagnostic: it is the visible footprint of where the recursion truncates, not evidence of failure to be rational.

Boundary-drawing — separate cannot-iterate from rationally-stops-short. The instrument lets the field draw a distinction the bare Nash benchmark erases. Reason FROM a finite-depth guess TO which of two cognitive states produced it: a player who cannot perform the iterated-dominance argument at all, versus a player who performs a few steps and then rationally halts because they expect others to halt too. The second reframes the "right" answer — it is not the equilibrium but a best response to one's belief about others' depth — so the move bounds when zero is even the correct target (only under common knowledge of rationality, which the game shows fails) and when the correct play is instead keyed to the population's expected reasoning depth.

Interventionist — vary a parameter and predict the shift. The protocol exposes two levers. Change the population (add game-theoretic training, professional financial expertise, or education) and predict that the distribution shifts toward lower guesses as reasoning depth rises — a prediction the instrument has confirmed, discriminating populations by exactly these factors and shifting student distributions downward after a course. Change the multiplier (from ⅔ to ½, or to any p) and the predicted per-level signatures move in a known way, letting the analyst test whether reasoning depth is task-invariant by reasoning FROM the new multiplier TO the new expected clusters and comparing. Either manipulation turns depth-of-reasoning into a dependent variable that can be pushed and measured.

Predictive / calibration — generate the data that fits cognitive-hierarchy models. The move that closes the loop is using the numerical signatures as the calibration target: from observed clusters, fit the depth-distribution parameters of cognitive-hierarchy models, then predict play in other recursive higher-order-belief tasks within the domain. Reasoning runs FROM measured truncation depth in this game TO a forecast of how the same population will behave wherever beliefs-about-beliefs determine the outcome — the game serving as the primary instrument whose readings the explanatory models are tuned against.

Knowledge Transfer

The ⅔ game is a measurement instrument — a named experimental protocol — not a causal mechanism, so the "mechanism within / metaphor beyond" frame does not apply; what transfers is an instrument, and it transfers literally wherever its precondition holds (a population of strategic participants who can submit a number under the winner-determination rule), with the boundary to mark being instrument-reach versus over-reading. Within behavioural game theory and experimental economics the protocol carries without translation across every population it can be run on: classrooms, online panels, lab studies with students and with professionals, and the famous newspaper-readership runs in the Financial Times and Spektrum der Wissenschaft. Because each level of iterated best-response leaves a distinct numerical signature (50, 33, 22, 14, …), the same instrument reads the modal depth of any population off its mean guess, and it has proven reliable as a discriminator — separating populations by education, game-theoretic training, and financial expertise, shifting students' distributions downward after a course — and as the calibration target whose readings cognitive-hierarchy models (Camerer-Ho-Chong) are tuned against. Varying the multiplier (½, p) or the bounds moves the per-level signatures in a known way, extending the instrument to test whether reasoning depth is task-invariant. All of this is the same protocol deployed wherever a strategic population exists; the home domain is the reach of the instrument itself.

The over-readings to guard against are the instrument-specific ones. First, treating a reading as a universal constant: the modal depth is a property of the population under these conditions, not a fixed human parameter — professionals reason deeper, trained students shift, stakes and framing move the number — so "people reason at level 2" over-reads a context-laden measurement as a species constant. Second, treating the gap from Nash as a verdict of irrationality rather than data: the cluster above zero is the visible footprint of finite-depth reasoning and of rationally stopping when one expects others to stop, not evidence that subjects cannot compute. Third, the purely verbal over-read: calling any situation with beliefs-about-beliefs "a ⅔ game" or "a beauty contest" borrows the name for a setting where no number is submitted and no level signature can be read — the instrument has not been deployed, only its imagery. Indeed the Keynesian beauty contest the game is the formal toy of, and the fund-manager pricing on others' forecasts of others' forecasts, are usually invoked metaphorically: they share the higher-order-belief skeleton, but they are not the instrument and yield no level reading, so the resemblance should be marked as analogy, not as the protocol traveling.

What genuinely generalizes is not "the ⅔ game" but the load-bearing patterns it operationalizes, and the cross-domain lesson should ride those parents. The finding that recursive higher-order belief truncates at finite depth, that common-knowledge predictions over-predict actual rationality, and that trading on others' beliefs about others' beliefs drives asset-price formation belong to higher_order_beliefs, bounded_rationality, common_knowledge, and the cognitive_hierarchy/level-k family — and to iterated_dominance, whose limit the game shows real populations never reach. When the insight is needed in financial trading, auction bidding, or common-value estimation, it is those general patterns that recur as co-instances and should carry it; the ⅔ game's specific contribution is to make the depth measurable by holding the recursion in a form where each layer leaves a countable trace. The instrument transfers literally across the populations where it can be run; its readings must not be over-read as constants or as failures; and the truncation it reveals travels onward under the names of the higher-order-belief and bounded-rationality primes it was built to probe — the division Structural Core vs. Domain Accent makes precise below.

Examples

Canonical

Rosemarie Nagel's experiment (1995) is the defining laboratory demonstration. Subjects each picked a number in [0, 100], winner closest to two-thirds of the average. The iterated-dominance argument is clean: the highest possible average is 100, so two-thirds of it is about 67, and no rational player picks above 67; but if no one exceeds 67, nothing above ~44 can win; iterating drives the unique Nash equilibrium to 0. Real subjects did not play 0. Their submissions clustered in spikes around 33 and 22 — precisely the level-1 response to a naive field guessing 50, and the level-2 response to a field guessing 33 — with the overall mean well above zero and only a handful choosing the equilibrium. The distribution's shape read directly as a census of how many steps of strategic reasoning the population performed.

Mapped back: The subjects picking numbers in [0,100] are the strategic population, and "closest to two-thirds of the mean" is the recursive winner rule. The chain to zero is the iterated-dominance endpoint that play never reaches; the spikes at 33 and 22 are the level-signature ladder made visible, and the mean sitting well above zero is the gap-from-Nash datum — the footprint of finite-depth reasoning, not error.

Applied / In Practice

The instrument has been run at scale on self-selected public populations. In 1997 the Financial Times, at Richard Thaler's instigation, invited readers to play the guess-two-thirds-of-the-average game for a prize. Roughly a thousand-plus entries came in from a sophisticated, largely financial readership. The average guess landed near 19, so two-thirds of it came out around 13 — and the winning number was 13. Notably, a cluster of entrants submitted 0 or 1 (players carrying the iterated-dominance argument all the way to equilibrium) and lost, because most of the field reasoned only a step or two deep. The readership's relatively low mean, compared with student samples, illustrated the instrument's discriminating power: a more strategically sophisticated population reasons visibly deeper.

Mapped back: The FT readers are the strategic population, and the contest applies the same recursive winner rule. The average near 19 and winning 13 are the modal-guess readout pinning the population's typical depth at roughly one to two steps. That this financial readership guessed lower than typical student samples is the parameter levers (population sophistication) at work — and the equilibrium-playing entrants who lost show why the iterated-dominance endpoint is the benchmark, not the winning action.

Structural Tensions

T1: Readable signature versus imposed model (the ladder measures depth only if level-k is assumed). The game's genius is that each step of iterated best-response maps to a distinct number — 50, 33, 22, 14 — so a modal guess seems to read a population's reasoning depth straight off the data. But that reading presupposes the level-k / cognitive-hierarchy generative model: a submission of 22 is labeled "level-2," yet the same number could come from a deep player expecting a shallow field, from risk-hedging toward the interior, from misreading the rule, or from noise. The distinct-signature clarity, in other words, is not model-neutral; it is diagnostic only once the level-k ladder is assumed to be what produces guesses. So the instrument that appears to measure reasoning depth objectively is in fact reading the data through the very model it is used to calibrate. The tension is that the countable trace which makes higher-order belief observable is a trace only under a specific theory of how the numbers are generated. Diagnostic: Is the modal guess being read as a level count under an assumed level-k model, or has it been checked against alternative processes (deep-player-shallow-belief, hedging, noise) that produce the same number?

T2: Depth measured versus reason for stopping underdetermined (capacity or rational belief?). The instrument claims to draw a distinction the bare Nash benchmark erases: a finite-depth guess could come from a player who cannot perform more iterations, or from one who performs a few and rationally halts because they expect others to halt too. But a single modal number pins the step-count while saying nothing about why the recursion truncated there — 22 is equally consistent with a level-2 ceiling and with a sophisticated player deliberately best-responding to a shallow field. The generous reading (rational finite depth) and the deflationary one (bounded capacity) are hard to separate from the number alone, and the generous reading risks unfalsifiability, since any guess is a best response to some belief about others. The tension is that the instrument reads depth cleanly yet leaves the interpretively decisive question — capacity versus rational stopping — underdetermined by its own output. Diagnostic: Does the evidence distinguish inability to iterate further from a rational choice to stop given beliefs about others' depth, or is a single modal guess being made to carry a distinction it cannot resolve?

T3: Discriminating power versus reified constant (a population reading, not a species parameter). The instrument's demonstrated value is that it discriminates — professionals reason deeper, trained students shift downward, the FT's financial readership guessed a mean near 19 and won at 13, well below student samples. That very sensitivity is what refutes any universal reading: the modal depth is a property of this population under these conditions, moved by expertise, stakes, and framing, not a fixed human constant. So the property that makes the game scientifically useful (it separates populations) is exactly the property that makes "people reason at level 2" a category error — a context-laden measurement reified into a species parameter. The tension is that the instrument's comparability across populations both grounds its power and invites the over-read that there is a single human depth to report. Diagnostic: Is the depth figure being used to compare matched populations, or is one population's context-laden reading being asserted as a general fact about how deeply humans reason?

T4: One-shot purity versus repetition-as-learning (measuring native depth requires naivety). The level-signature reading is clean only in a genuinely one-shot game against a naive field: it captures how deep a population reasons before feedback teaches it anything. But the game is frequently run repeatedly, and under repetition the winning number falls toward zero as players learn from the previous round's outcome — a convergence driven by feedback and adaptation, not by native strategic depth. So repetition contaminates the very quantity the instrument exists to measure: after a few rounds the low guesses reflect learning about this game, not the population's baseline reasoning depth. The tension is that the instrument's readout is valid precisely when players are naive, yet the natural way to sharpen or demonstrate it (repeat the game, let them learn) destroys the naivety that made the first reading a measurement of native depth. Diagnostic: Is the guess distribution being read from a naive one-shot play, or from a repeated game where falling numbers reflect learning and convergence rather than native reasoning depth?

T5: Autonomy versus reduction (a laboratory protocol or the higher-order-belief primes it operationalizes). The guess-⅔ game is not a causal mechanism but a named experimental protocol, so it transfers literally — not by analogy — wherever a strategic population can submit numbers under the winner rule: classrooms, online panels, the Financial Times and Spektrum runs. Its proprietary apparatus (the recursive winner rule, the level-signature ladder, the modal-guess readout, the multiplier lever) is what carries there. But reduction bites at both edges: calling any beliefs-about-beliefs situation "a ⅔ game" or "a beauty contest" — a fund manager pricing on others' forecasts of others' forecasts — borrows the imagery for a setting where no number is submitted and no level reading exists, which is metaphor; and the findings generalize not as "the ⅔ game" but as instances of higher_order_beliefs, bounded_rationality, common_knowledge, the cognitive_hierarchy/level-k family, and iterated_dominance (whose limit real play never reaches). The tension is between an instrument that travels literally where runnable and the higher-order-belief primes that carry its lessons everywhere else. Diagnostic: Resolve toward the primes (higher_order_beliefs, bounded_rationality, cognitive_hierarchy, iterated_dominance) when the point is truncated recursive belief in the field; toward the named protocol only when the actual instrument — the winner rule and its level signatures — is being run on a population.

Structural–Framed Character

The guess-⅔-of-the-average game occupies an unusual position — best read as mixed, with the qualification the entry insists on: it is not a mechanism or a phenomenon but a measurement instrument, a designed experimental protocol that reads a hidden cognitive quantity off an observable number. On evaluative_weight it reads structural: the instrument is evaluatively inert — it measures depth of strategic reasoning and convicts nothing, and the entry is emphatic that the gap from Nash is "the visible footprint of finite-depth reasoning, not evidence that players cannot reason," so it renders a reading, not a verdict. That neutrality is its main structural mark. On the other four criteria it reads framed, which holds it at mixed. On human_practice_bound it reads strongly framed: the game exists only as a set-up experiment — a strategic population submitting numbers under a winner rule — a contrivance constituted entirely by the practice of running it, with no observer-free existence at all. Institutional_origin is likewise framed: it is a designed protocol (Moulin 1986, Nagel 1995), an artifact of behavioral game theory built to elicit a reading. On vocab_travels it reads framed: the recursive winner rule, the iterated-dominance endpoint at zero, the level-signature ladder (50, 33, 22, 14), cognitive-hierarchy calibration are pinned to game-theoretic experiment. And on import_vs_recognize the transfer is bimodal exactly as the entry documents — the protocol transfers literally (recognition) wherever a strategic population can be run (classrooms, online panels, the FT and Spektrum surveys), but "a ⅔ game" or "a beauty contest" invoked for any beliefs-about-beliefs setting where no number is submitted is import-by-imagery, metaphor, not the instrument deployed.

The instrument's structural anchor is that what it reads is real: finite-depth recursive reasoning is a genuine cognitive phenomenon that exists in people's heads independent of the experiment, and the level-signature ladder is a real mathematical mapping from step-count to number. But that does not lift the entry toward the structural pole, because the phenomenon belongs to the parents while the game is the artifact that measures it. The portable structural skeleton is a composition of the higher-order-belief primes it operationalizes — higher_order_beliefs, bounded_rationality, common_knowledge, the cognitive_hierarchy/level-k family, and iterated_dominance (whose limit real play never reaches). That composition genuinely recurs across financial trading, auction bidding, and common-value estimation, but it is exactly what the game operationalizes from its parents, not what makes "the ⅔ game" itself travel: the cross-domain lesson (recursive belief truncates at finite depth; common-knowledge predictions over-predict rationality) belongs to those primes, while the winner rule, the level-signature ladder, and the modal-guess readout stay home as the instrument's own apparatus. Its character: an evaluatively neutral but wholly practice-bound experimental protocol that reads a real cognitive quantity off a number — mixed, structural in its neutrality and in the higher-order-belief phenomenon it measures, framed in the designed, protocol-bound contrivance that is actually its own.

Structural Core vs. Domain Accent

This section decides why the guess-⅔-of-the-average game is a domain-specific abstraction and not a prime — and the case is distinctive because the game is not a mechanism but a measurement instrument, so its "structural core" is the cognitive phenomenon it reads, not anything the protocol itself contains.

What is skeletal (could lift toward a cross-domain prime). Strip the experimental protocol and what survives is not a mechanism but the phenomenon it measures: recursive higher-order belief — reasoning about what others believe about what others believe — truncates at a finite depth, so predictions built on common knowledge of full rationality over-predict how deeply agents actually reason. The portable pieces are abstract and are themselves primes — beliefs about beliefs (higher_order_beliefs), a finite reasoning budget (bounded_rationality), the common-knowledge premise (common_knowledge) whose failure the truncation exposes, the ladder of finite-depth types (cognitive_hierarchy/level-k), and the iterated-dominance chain (iterated_dominance) whose limit real play never reaches. That composition genuinely recurs — financial trading, auction bidding, common-value estimation. But it is what the game operationalizes, not what the game is: the protocol's own contribution is to render the recursion in a form where each layer leaves a countable trace, not to supply the phenomenon.

What is domain-bound. Everything specific to the object is game-theoretic-experiment furniture. The set-up is not generic — it is a strategic population submitting a number in [0,100] under the recursive winner rule (closest to ⅔ of the mean). The reading apparatus is a worked mapping — the level-signature ladder (level-0 ≈ 50, level-1 ≈ 33, level-2 ≈ 22, level-3 ≈ 14) — and its output is the modal-guess readout pinning a population's modal depth, calibrated against cognitive-hierarchy models. Its levers (the multiplier ⅔ → ½ → p, the population) and its worked runs (Nagel 1995, the Financial Times contest won at 13) are all designed experiments. The decisive test: remove the population submitting numbers under the winner rule and there is no ⅔ game left — only the underlying higher-order-belief phenomenon, which a fund manager pricing on others' forecasts exhibits with no number to submit and no level signature to read. What remains after stripping the protocol is the bare truncated-recursion phenomenon, which belongs to the parents.

Why this does not clear the prime bar. A prime's vocabulary travels and its transfer is recognition of the same mechanism. The game's reach is unusual: as an instrument it transfers literally — not by analogy — wherever a strategic population can be run under the winner rule (classrooms, online panels, the FT and Spektrum surveys), because the same level-signature ladder reads any such population's depth off its mean guess. But that literal reach is the reach of the instrument, one experimental protocol deployed across populations, not cross-substrate structural travel. Beyond the runnable set-up the game does not travel at all: calling any beliefs-about-beliefs situation "a ⅔ game" or "a beauty contest" borrows the imagery for a setting where no number is submitted and no reading exists — metaphor, not the instrument. And when the substantive lesson is wanted elsewhere — recursive belief truncates at finite depth; common-knowledge predictions over-predict rationality — it is carried by the phenomenon's parents (higher_order_beliefs, bounded_rationality, common_knowledge, cognitive_hierarchy, iterated_dominance), which recur as genuine co-instances in trading and auctions. The cross-domain reach belongs to those primes; "the ⅔ game," as named, is the protocol whose winner rule, level-signature ladder, and modal-guess readout are the instrument's own apparatus and should stay home — with the added caution that its readings are population-and-condition-specific measurements, never universal human constants.

Relationships to Other Abstractions

Local relationship map for Guess 2/3 of the AverageParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Guess 2/3 ofthe AverageDOMAINDomain-specific abstraction: Beauty Contest Game — is a kind ofBeautyContest GameDOMAIN

Current abstraction Guess 2/3 of the Average Domain-specific

Parents (1) — more general patterns this builds on

  • Guess ⅔ of the Average is a kind of Beauty Contest Game Domain-specific

    Guess ⅔ of the Average is the Beauty Contest Game with p fixed to two thirds, the action interval fixed to 0–100, and its canonical level ladder exposed.

Hierarchy paths (9) — routes to 7 parentless roots

Not to Be Confused With

  • The Keynesian beauty contest. The informal metaphor the ⅔ game is the formal toy of — Keynes's image of investors picking not the prettiest face but the face they expect others to expect others to pick, and the fund manager pricing on others' forecasts of others' forecasts. It shares the higher-order-belief skeleton but submits no number and yields no level signature, so it is analogy, not the instrument deployed. Tell: is a strategic population actually submitting numbers under the winner rule (the game), or is "beauty contest" borrowed as imagery for a beliefs-about-beliefs setting (the metaphor)?
  • The cognitive-hierarchy / level-k model. The generative theory — agents best-respond to a finite-depth distribution of shallower players — that the game is used to calibrate. The model is the explanation; the ⅔ game is the measurement protocol whose numerical clusters supply the model's data. Conflating them treats the reading (a modal guess) as the theory that interprets it. Tell: is this the account of why depth is finite (the model), or the experimental apparatus that reads depth off a number (the game)?
  • The Nash equilibrium of zero. The game's benchmark, not a prediction of play or a recommended action: iterated dominance drives the unique equilibrium to zero, yet real populations cluster well above it, and submitting zero loses against a field around 22. The equilibrium is the reference the observed gap departs from, not what the game predicts people do. Tell: is zero being cited as the common-knowledge-of-rationality solution (the benchmark) or as the winning play (which it is not)?
  • Iterated elimination of dominated strategies. The solution procedure that defines the endpoint — chaining "no number above the current ceiling can win" down to zero — a general game-theoretic method the game happens to make vivid. The ⅔ game's contribution is not this procedure but the finding that real play truncates it after a step or two, leaving a countable trace. Tell: is the point the abstract dominance-chaining argument (the procedure) or the empirical depth at which a population stops performing it (the game's reading)?
  • The centipede game. A neighbor experimental game where iterated reasoning — here backward induction — likewise predicts an equilibrium (defect immediately) that real players systematically fail to reach, cooperating for several rounds. Both dramatize common-knowledge-of-rationality over-predicting actual play, but the centipede reads deviation off a stopping node in a sequential tree, not off a number keyed to a level-signature ladder. Tell: is depth being read from a single simultaneous numeric submission (⅔ game) or from how far players proceed in a sequential take-or-pass tree (centipede)?
  • The higher-order-belief primes it operationalizes (higher_order_beliefs, bounded_rationality, common_knowledge, the cognitive_hierarchy/level-k family, iterated_dominance). The substrate-neutral umbrella whose lessons — recursive belief truncates at finite depth; common-knowledge predictions over-predict rationality — genuinely recur in trading, auctions, and common-value estimation. Tell: strip the winner rule, the level-signature ladder, and the modal-guess readout and what remains is the truncated-recursion phenomenon — the parents, treated more fully elsewhere, not "the ⅔ game."

Neighborhood in Abstraction Space

Guess ⅔ of the Average sits in a crowded region of the domain-specific corpus (0th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Strategic Interaction & Game Theory (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12