Skip to content

Worse-than-Average Effect

The reversal of self-enhancement on hard, expert-dominated tasks, where most people rate themselves below average because self-ratings anchor on one's own absolute perceived skill and ignore the distribution of others — the sign of the bias set by perceived task difficulty.

Core Idea

The worse-than-average effect is the empirically observed reversal of self-enhancement bias: on tasks perceived as difficult, unfamiliar, or heavily skewed toward expertise — juggling, chess, programming, complex surgical procedures — people on average rate their own ability as below the population average, even though by definition the majority cannot be below the median. The effect is the direct counterpart to the better-than-average effect on easy tasks, and both arise from the same underlying mechanism: Kruger (1999) demonstrated that self-assessments are anchored primarily on one's own absolute perceived skill level, rather than on an accurate model of the distribution of others' skills. On an easy task where nearly everyone has some competence, anchoring on absolute self-assessed skill places most respondents above what they estimate the average to be; on a hard task dominated by a few clear experts, the same anchoring places most respondents below it. The distributional incoherence in both directions — more than 50% cannot simultaneously be above or below the median — is the diagnostic signature that the self-assessment process is not tracking the actual comparative distribution.

Structural Signature

Sig role-phrases:

  • the self-assessor — the person rating their own ability against others
  • the difficulty-varying task — a task whose perceived difficulty is the one moderator that sets the sign (hard/expert-dominated here, vs. the easy case for the better-than-average twin)
  • the reference class — the imagined population distribution the self-rating is supposed to be compared against
  • the absolute-skill anchoring — the load-bearing mechanism: self-ratings anchor on one's own absolute perceived skill and largely ignore the distribution of others
  • the missing-distribution stage — the localised defect: self-judgment splits into a roughly-accurate absolute-self read and an often-absent model of where others stand, and the bias lives in the second
  • the below-average reversal — the signature on hard tasks: most people rate themselves below average, the obverse of the better-than-average effect
  • the population-incoherence diagnostic — the tell that the reference class was never consulted: more than half cannot truly be below the median
  • the difficulty cross-over — the unifying prediction: reframe a task harder or easier and the sign flips, the two effects one curve indexed by difficulty

What It Is Not

  • Not accurate modesty. A sample in which most people rate themselves below average is not realistic self-knowledge: by definition the majority cannot be below the median, so the pattern is the population-incoherence signature of an absolute-anchoring process, not a population that genuinely "sells itself short." The below-average reversal is an artifact of how the comparison is made, not a true read of comparative standing.
  • Not a fixed, unidirectional self-evaluation bias. Self-assessment bias does not always run one way: its direction is set by perceived task difficulty. Easy tasks yield over-rating (the better-than-average effect), hard tasks yield under-rating (this effect), and the two are one cross-over curve — so "people flatter themselves" is only half the picture, and the sign is predictable from the task.
  • Not a defect in self-knowledge of one's own ability. The bias lives in the missing model of where others stand, not in the read of one's own competence, which is often roughly accurate. So the operative question is not "are people miscalibrated about themselves?" but "did the judgment consult a reference class at all?" — which is why supplying distributional information about others is the repair.
  • Not Dunning–Kruger, impostor syndrome, or regression to the mean. It is the average direction of self-rating set by task difficulty across all competence levels — distinct from Dunning–Kruger's skill-quartile claim that the least competent overestimate most, from impostor syndrome's felt sense of fraudulence, and from regression to the mean's statistical artifact of noisy repeated measurement. These share the self-assessment neighbourhood but make different claims.
  • Not the substrate-free anchoring pattern. Stripped of the self-rating paradigm and the population-incoherence diagnostic, the residue — a judgment anchored on the focal object's absolute properties that misses the comparative distribution — simply is anchoring_bias, of which this effect and its better-than-average twin are the difficulty-modulated obverse halves. The cross-domain lesson belongs to the parent; outside self-assessment, "worse-than-average" is just reference-class-neglect anchoring under a borrowed name.

Scope of Application

The worse-than-average effect lives across social and cognitive psychology and behavioural economics wherever a self-assessing agent rates itself against an imagined population on a difficulty-varying task; its reach is bounded to that self-assessment substrate (the underlying reference-class-neglect anchoring travels further under its anchoring_bias parent, and the within-family neighbours Dunning–Kruger and impostor syndrome make different claims).

  • Social psychology of self-assessment — a canonical exhibit in the biased-self-evaluation literature, the difficulty-modulated obverse of the better-than-average effect.
  • Educational assessment — systematic under-rating on objectively hard subjects (advanced calculus, organic chemistry), with downstream effects on course-taking and persistence.
  • Career self-selection — under-application by competent candidates to skilled-technical domains where the imagined distribution is expert-dominated.
  • Comparative health-risk perception — below-average self-ratings on rare, high-skill protective behaviours (e.g. recognising skin cancer).
  • Survey methodology — a known comparative-judgment artifact requiring researcher correction.

Clarity

Recognizing the worse-than-average effect corrects the field's once-tidy story that self-assessment bias runs in a single direction — that "people flatter themselves." Naming the below-average reversal makes visible that the direction of the bias is not a fixed feature of self-evaluation but a function of perceived task difficulty: easy tasks yield over-rating, hard tasks yield under-rating. That reframing dissolves an apparent contradiction in the literature — over-confidence in some studies, under-confidence in others — that looked like inconsistent findings until the difficulty axis was made the organizing variable. Once it is, the two effects are seen as a single cross-over rather than two unrelated curiosities, and a researcher can predict the sign of the bias from a task's difficulty instead of measuring it case by case.

The deeper clarity is locating the defect precisely. The label points the analyst at the anchoring step: self-ratings track one's own absolute perceived skill and largely ignore the comparative distribution of others. This separates two things comparative-judgment surveys otherwise blur — a person's read of their own competence (often accurate) from their model of where others stand (often absent) — and shows the bias lives in the second, not the first. That makes the population-incoherence pattern (more than half claiming to be below the median) legible as a diagnostic that the comparative distribution was never consulted, and it sharpens the operative question for debiasing: not "are people miscalibrated about themselves?" but "did the judgment incorporate a reference class at all?" — predicting that supplying distributional information about others should pull both the easy-task and hard-task biases toward coherence.

Manages Complexity

The self-assessment literature, read superficially, is a contradiction: some studies find people grossly over-rate themselves (driving, getting along with others, common sense), others find them under-rate themselves (chess, surgery, programming), and the two bodies of results seem to license opposite generalizations about human self-knowledge — "people flatter themselves" versus "people sell themselves short." Catalogued effect by effect, every comparative-judgment domain becomes its own finding, and the sign of the bias has to be measured anew each time. The worse-than-average effect, taken together with its better-than-average twin, compresses that whole field onto a single mechanism with one continuous moderator the analyst can track: self-ratings anchor on one's absolute perceived skill and largely ignore the distribution of others, and the only thing that varies across tasks is perceived difficulty. Given a task's position on that one difficulty axis, the sign of the bias reads off directly — easy tasks (nearly everyone competent, the imagined distribution low-variance and high) place most self-ratings above the estimated average; hard tasks (a few clear experts dominating the imagined distribution) place most self-ratings below it. What looked like two unrelated curiosities, or an inconsistent literature, collapses to one cross-over curve indexed by difficulty, and the population-incoherence signature (more than half claiming to be below the median) is the same diagnostic in both regimes that the comparative distribution was never consulted.

The deeper compression is in where the defect is located, which turns a vague "people are miscalibrated" into a structured two-stage account the analyst can run. Self-judgment is split into an absolute-self read (often roughly accurate) and a distribution-of-others model (often simply absent), and the entire bias is assigned to the second stage. That decomposition lets a researcher do three things from a small set of tracked quantities — perceived task difficulty, whether a reference class was supplied — without re-deriving each domain: predict the sign of the bias before measuring it; predict that the bias flips as a task is reframed harder or easier (the cross-over); and predict the single intervention that should pull both the easy-task and hard-task biases toward coherence, namely supplying distributional information about others. The sprawl of direction-specific self-assessment findings reduces to one anchoring mechanism, one difficulty moderator, and a localized second-stage failure whose remedy is the same wherever the effect appears.

Abstract Reasoning

The worse-than-average effect licenses reasoning that treats comparative self-judgment as anchoring on absolute perceived skill while ignoring the distribution of others, with perceived task difficulty the one moderator that sets the sign — so the researcher reasons from a task's difficulty to the direction of the bias, and from a population-incoherence pattern back to a specific stage of the judgment that failed.

Diagnostic (read the sign from difficulty; read the failed stage from incoherence). The defining inference goes from a comparative-judgment result to the mechanism behind it. A sample in which most people rate themselves below average is read not as accurate modesty but as the absolute-anchoring process applied to a hard, expert-dominated task; the same process on an easy task is read as producing the better-than-average pattern — so the bias's direction is diagnosed from the task's difficulty rather than from any property of self-evaluation in general. The population-incoherence signature — more than half claiming to be below (or above) the median — is the diagnostic that the comparative distribution was never consulted, and it points the analyst at the second of two stages: self-judgment splits into an absolute-self read (often roughly accurate) and a distribution-of-others model (often simply absent), and the bias is localized to the latter. The inference runs result + task difficulty → absolute-anchoring on a hard task (and incoherence → "the reference class was skipped"), never result → "people sell themselves short."

Interventionist (reframe difficulty or supply a reference class, predict the shift). Because the sign is set by perceived difficulty and the defect by a missing distribution, both are levers with forecasts. Reframe a task as harder and average self-ratings are predicted to shift downward, crossing over from the better-than-average regime toward the worse-than-average one; reframe it easier and they shift upward — the cross-over is a manipulation, not just an observation. Supply distributional information about how others actually perform and the prediction is that both the easy-task and hard-task biases pull toward coherence, because the intervention repairs exactly the second stage the bias lives in. Each manipulation pairs a change in difficulty or in reference-class availability with a predicted, signed change in the comparative judgment.

Boundary-drawing (the defect is the distribution stage, not the self-read; difficulty is the axis). The concept draws two lines. It separates a person's read of their own competence (often accurate) from their model of where others stand (often absent), and assigns the bias to the second — so "are people miscalibrated about themselves?" is the wrong question and "did the judgment incorporate a reference class at all?" is the right one. And it makes difficulty the organizing axis that bounds where each direction of bias is expected: easy tasks over the cross-over yield over-rating, hard tasks under it yield under-rating, and the two are one curve rather than two unrelated effects. Those boundaries also separate the effect from neighbors — it is the average direction of self-rating by task difficulty, not the confidence-competence miscalibration of the least skilled, nor a felt sense of fraudulence, nor a statistical regression artifact.

Predictive / ordering. From a task's position on the single difficulty axis the analyst forecasts the sign of the bias before measuring it, and forecasts that the same population flips as the task is reframed across the cross-over; from whether a reference class was supplied, the analyst forecasts whether the incoherence will persist or resolve. Direction, the cross-over, and the debiasing response all read off difficulty plus reference-class availability, with the absolute self-read held fixed as the part the bias does not touch.

Knowledge Transfer

Within social and cognitive psychology and behavioural economics the effect transfers as mechanism, because the anchoring account (self-ratings track absolute perceived skill and ignore the distribution of others) and its single moderator (perceived task difficulty) apply unchanged across the cases. It is a canonical exhibit in the biased-self-evaluation literature; in educational assessment it predicts systematic under-rating on objectively hard subjects, with downstream effects on course-taking and persistence; in career self-selection it predicts under-application by competent candidates to skilled-technical domains; in comparative health-risk research it predicts below-average self-ratings on rare, high-skill protective behaviours; and in survey methodology it is a known comparative-judgment artifact requiring correction. The vocabulary (better/worse-than-average cross-over, absolute-skill anchoring, reference class, population incoherence, difficulty moderator) and the two-part prediction (read the sign off difficulty; repair the bias by supplying a distribution) carry intact across that cluster because the substrate is constant: a self-assessing agent rating itself against an imagined population distribution.

Beyond that self-assessment substrate the entry is a clean shared abstract mechanism (B), and the parent is already in the catalogue. The genuinely substrate-spanning structure is not the worse-than-average effect but the anchoring pattern it instantiates — a judgment anchored on absolute properties of the focal object misses the comparative distribution — which the catalogue carries as anchoring_bias, and of which this effect and its better-than-average twin are the difficulty-modulated obverse halves (with dunning_kruger_effect the neighbouring miscalibration finding). So when the cross-domain lesson is wanted — "rating something by its own absolute attributes, without consulting the reference class, produces distributionally incoherent comparative judgments" — it should be carried by anchoring_bias (and the difficulty-modulated self-assessment family), not by "worse-than-average effect," whose distinctive cargo (the self-rating paradigm, the population-incoherence diagnostic, the specific better/worse cross-over indexed by task difficulty) is self-assessment furniture that does not travel. The effect's interventions — reframe difficulty to flip the sign, supply distributional information to restore coherence — are likewise specific to comparative self-judgment and do not port; the portable move is the anchoring-on-absolute-property insight the parent already states.

The boundary to mark is twofold. First, the cross-domain reach beyond psychology and behavioural economics is genuinely shallow: the effect needs a self-assessing agent with a reference class, so invoking "worse-than-average" outside that setting borrows the label for what is, structurally, plain reference-class-neglect anchoring — to be named at the parent, not as this effect. Second, the within-family neighbours must not be conflated with it: Dunning–Kruger (the least competent overestimate most, a skill-quartile claim) and impostor syndrome (a felt sense of fraudulence) are different claims from the worse-than-average effect (an average direction of self-rating set by task difficulty, across all competence levels), even though all sit in the self-assessment literature. The clean boundary, then: literal transfer of the worse-than-average effect across comparative self-assessment contexts wherever an agent rates itself against a population on a difficulty-varying task; and beyond that, the reach belongs to anchoring_bias and the difficulty-modulated self-assessment family it instantiates, not to the named effect. (See Structural Core vs. Domain Accent.)

Examples

Canonical

Justin Kruger's 1999 studies ("Lake Wobegon be gone!") supplied the defining demonstration. Participants rated their ability relative to their peers across a range of tasks that differed in perceived difficulty. On easy, familiar tasks — using a computer, driving — people showed the usual better-than-average pattern, placing themselves above the median. But on tasks perceived as difficult and expertise-dominated — computer programming, chess, juggling — the pattern reversed: participants on average rated themselves below the peer median, a statistical impossibility for more than half a population. Critically, Kruger showed why: people's comparative self-ratings tracked their own absolute sense of skill on the task far more strongly than any accurate estimate of how others performed. On hard tasks people feel unskilled in absolute terms and therefore rank themselves low, never properly consulting the distribution of everyone else, who mostly feel unskilled too.

Mapped back: The rating participants are the self-assessors, and varying tasks from easy to hard is the difficulty-varying task whose perceived difficulty sets the sign. Ratings driven by absolute felt skill are the absolute-skill anchoring, and the failure to model peers' performance is the missing-distribution stage. Most people placing themselves below the median on hard tasks is the below-average reversal, and its impossibility is the population-incoherence diagnostic that the reference class was never consulted.

Applied / In Practice

Education researchers use the effect to explain and address why capable students abandon hard subjects. In objectively difficult fields — advanced mathematics, organic chemistry, computer science — students frequently rate their own ability below their peers because the material feels hard in absolute terms, and this under-rating predicts dropping the course, switching majors, or not pursuing the field even when the student is performing adequately relative to classmates. The effect is a live concern in STEM-persistence and equity work, where talented students self-select out. The theory also prescribes the remedy: because the defect is a missing model of where others stand, supplying honest distributional feedback — "most students find this problem set hard; here is the class performance distribution" — is predicted to pull self-ratings toward coherence and reduce the unwarranted attrition, which normative-feedback and belonging interventions attempt in practice.

Mapped back: The struggling-but-adequate student is the self-assessor on a hard difficulty-varying task; feeling the material is hard in absolute terms is the absolute-skill anchoring producing the below-average reversal that drives them out. The intervention — showing the class performance distribution — directly repairs the missing-distribution stage, supplying the reference class the biased judgment skipped.

Structural Tensions

T1: Direction set by difficulty versus a fixed self-enhancement bias (the cross-over). The field's tidy story — "people flatter themselves" — is only half the picture: the direction of self-assessment bias is not a fixed feature of the ego but a function of perceived task difficulty, over-rating on easy tasks and under-rating on hard ones, one continuous cross-over curve. The tension is that this dissolves an apparent contradiction (over-confidence in some studies, under-confidence in others) only by making difficulty the organizing variable, which means neither "people are over-confident" nor "people sell themselves short" is a stable generalization — each is a regime of one curve. Any claim about the sign of self-assessment bias is under-specified until the task's difficulty is fixed, and the intuitive default (self-enhancement) is exactly wrong for hard, expert-dominated tasks. Diagnostic: Is the task here easy (predict over-rating) or hard and expert-dominated (predict under-rating) — and is a fixed-direction self-enhancement assumption being applied across the cross-over?

T2: Missing distribution-of-others versus accurate self-read (where the defect lives). The bias is localized not in a person's read of their own competence — which is often roughly accurate — but in the second stage, the model of where others stand, which is frequently simply absent. This is the construct's sharp contribution: it splits comparative self-judgment into an absolute-self read and a distribution-of-others model and assigns the entire bias to the latter. The tension is that the two stages are fused in the single comparative rating a survey elicits ("rate yourself relative to peers"), so the accurate self-knowledge and the missing reference class produce one number that looks like a failure of self-knowledge when it is a failure of population-modeling. Treating the below-average rating as miscalibration about oneself points the remedy at the wrong stage. Diagnostic: Is the error in the person's read of their own ability, or in an absent model of how everyone else performs — and is the intervention aimed at the stage that actually failed?

T3: Incoherent artifact versus accurate modesty (a below-average that cannot be true). When most people rate themselves below the median, it looks like a population that realistically recognizes its limitations — but more than half cannot truly be below the median, so the pattern is a distributional-incoherence signature, not accurate modesty. The tension is that the surface reading (humility, realism) is the opposite of the structural reading (an anchoring artifact that never consulted the distribution), and the two are hard to tell apart at the level of any single rating, since a genuinely below-average person and an incoherent self-anchorer give the same answer. Only the population-level impossibility exposes the artifact. Crediting the below-average pattern as honest self-knowledge misreads a diagnostic of skipped comparison as a virtue. Diagnostic: Is this below-average self-rating genuine comparative accuracy, or the population-incoherence signature (more than half below the median) that the reference class was never consulted?

T4: The effect versus its self-assessment neighbours (different claims in one literature). The worse-than-average effect is an average direction of self-rating set by task difficulty, across all competence levels — and it is routinely conflated with neighbours that make different claims: Dunning-Kruger (a skill-quartile claim that the least competent overestimate most), impostor syndrome (a felt sense of fraudulence), and regression to the mean (a statistical artifact of noisy repeated measurement). The tension is that all four sit in the self-assessment neighbourhood and can co-occur in the same discouraged student, so an observed under-rating invites attribution to whichever neighbour is most familiar, blurring distinct mechanisms with distinct remedies. Treating the worse-than-average effect as "just Dunning-Kruger" imports a competence-quartile story where the claim is difficulty-indexed and level-general. Diagnostic: Is the phenomenon here difficulty-set average direction (worse-than-average), a competence-quartile miscalibration (Dunning-Kruger), a felt fraudulence (impostor), or a measurement artifact (regression)?

T5: Coherence restored versus behaviour unmoved (what the reference-class remedy does and does not fix). Because the defect is the missing distribution, supplying honest distributional feedback is predicted to pull both easy- and hard-task biases toward coherence — a clean, mechanism-matched remedy. But the downstream behaviour the effect is invoked to explain (capable students dropping hard subjects) is driven not only by the incoherent comparative rating but by the absolute felt difficulty of the material, which the anchoring account holds is often accurate and which distributional feedback does not touch. The tension is that restoring comparative coherence corrects the miscoded judgment without necessarily removing the absolute discouragement that co-drives the attrition, so the remedy targets the biased stage precisely yet may under-deliver on the behaviour, which is why belonging and normative-feedback interventions are needed alongside. Diagnostic: Is the goal restoring comparative coherence (which distributional feedback fixes) or changing the behaviour (which also turns on absolute felt difficulty the reference class does not address)?

T6: Autonomy versus reduction (a self-assessment effect or an instance of anchoring bias). The worse-than-average effect is a named finding with proprietary apparatus — the self-rating paradigm, the population-incoherence diagnostic, the better/worse cross-over indexed by task difficulty — and within social and cognitive psychology it transfers as mechanism intact. But strip that apparatus and the residue — a judgment anchored on the focal object's absolute properties that misses the comparative distribution — simply is anchoring_bias, of which this effect and its better-than-average twin are the difficulty-modulated obverse halves. Beyond a self-assessing agent with a reference class, "worse-than-average" is just reference-class-neglect anchoring under a borrowed name, and the effect's interventions (reframe difficulty, supply a distribution) are self-judgment-specific and do not port. The tension is between a self-assessment finding that earns its own name and the recognition that its cross-domain reach belongs to the anchoring parent. Diagnostic: Resolve toward anchoring_bias when the lesson is absolute-property anchoring that skips the reference class in any judgment; toward the named effect when diagnosing the direction of self-ratings on a difficulty-varying task.

Structural–Framed Character

The worse-than-average effect sits at the mixed position on the structural–framed spectrum — a genuine cognitive mechanism, which pulls toward structure, but a mind-bound bias carrying an implicit coherence norm and reducing almost entirely to its parent, which holds it back. The criteria split. On evaluative_weight it is mildly framed: the effect is defined against a normative benchmark — distributional coherence — and its diagnostic signature is literally an incoherence (more than half claiming to be below the median), so calling a pattern "worse-than-average" flags a departure from what a well-calibrated comparison would yield, a value-laden reading absent from a pure regularity like Weber's law; but it is a mechanism (absolute-skill anchoring), not a verdict on a person's character, so the tint is faint. On institutional_origin it points structural: the effect is a fact of how self-assessment anchors on absolute perceived skill and skips the reference-class distribution, discovered (Kruger), not an artifact of any convention. And it is not human-practice-bound: a self-assessing agent under-rates on hard tasks whether or not any psychologist runs a comparative-judgment survey — remove every experimenter and the absolute-anchoring process still misfires, so nothing dissolves when the scholarly practice is withdrawn.

What pulls it toward the framed side of mixed is the combination of a narrow mind-only substrate and near-total reducibility. On vocab_travels it fails beyond self-assessment: absolute-skill anchoring, reference class, population incoherence, difficulty cross-over are pinned to a self-rating agent, and outside that setting "worse-than-average" is just reference-class-neglect anchoring under a borrowed name. On import_vs_recognize it is recognition within psychology and behavioural economics (the same mechanism and difficulty moderator carry across education, career self-selection, health-risk perception) but beyond it the effect is not imported — the parent is. And its portable content is thin: stripped of the self-rating paradigm and the incoherence diagnostic, the residue — a judgment anchored on the focal object's absolute properties that misses the comparative distribution — simply is anchoring_bias.

The portable structural skeleton is anchoring on a focal object's absolute attributes while neglecting the comparative distribution, yielding distributionally incoherent comparison — and this is what the effect instantiates from its parent prime anchoring_bias (with its better-than-average twin as the difficulty-modulated obverse half, and dunning_kruger_effect a neighbouring finding), not what makes "worse-than-average effect" itself travel. The cross-domain reach belongs to anchoring_bias; the self-rating paradigm, the population-incoherence diagnostic, and the difficulty cross-over stay home. Its character: a real, discovered self-assessment mechanism, faintly evaluative through its coherence benchmark and mind-bound in substrate, structural only in the absolute-anchoring-skips-the-distribution skeleton it instantiates from anchoring_bias — mixed, held off the structural side by a mind-only substrate, an implicit coherence norm, and near-total reducibility to its parent.

Structural Core vs. Domain Accent

This section decides why the worse-than-average effect is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — there is no separate section for that.

What is skeletal (could lift toward a cross-domain prime). Strip the self-rating paradigm and a thin relational structure survives: a comparative judgment anchored on the focal object's absolute attributes neglects the comparative distribution, yielding a distributionally incoherent verdict. The portable pieces are abstract — a focal object with absolute properties, a reference distribution the comparison should consult, an anchoring step that reads the absolute properties, and a missing step that skips the distribution. That skeleton is genuinely substrate-portable, which is why the entry attributes it to the catalog prime the effect instantiates: anchoring_bias, of which this effect and its better-than-average twin are the difficulty-modulated obverse halves (with dunning_kruger_effect a neighbouring finding). That anchor-on-absolute-skip-the-distribution core is what the worse-than-average effect shares with any absolute-property anchoring that neglects a reference class — not what makes it the worse-than-average effect.

What is domain-bound. The distinctive content is self-assessment furniture and none of it survives extraction intact: the self-assessor rating their own ability; the difficulty-varying task whose perceived difficulty is the one moderator that flips the sign; the absolute-skill anchoring on one's own felt competence; the missing-distribution stage (the split between a roughly-accurate absolute-self read and an absent model of others); the below-average reversal as the signature on hard tasks; the population-incoherence diagnostic (more than half cannot truly be below the median); and the difficulty cross-over that unites this effect with the better-than-average twin as one curve. These are the worked vocabulary, the instruments (comparative-judgment surveys), and the empirical cases the field studies — Kruger's easy/hard tasks, STEM-persistence attrition, career self-selection, health-risk perception. The decisive test: remove the self-assessing agent with a reference class — invoke "worse-than-average" for some judgment with no self-rating and no imagined population — and it is no longer this effect but plain reference-class-neglect anchoring under a borrowed name, which belongs to the parent. The effect is constituted by the self-assessment substrate the prime bar asks it to shed.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. The worse-than-average effect's transfer is bimodal. Within social and cognitive psychology and behavioural economics it travels intact as mechanism — the anchoring account and its single difficulty moderator carry unchanged across educational assessment, career self-selection, comparative health-risk perception, and survey methodology, which are one mechanism, not analogies, and the sign of the bias is predictable from the task in each. Beyond self-assessment the reach is genuinely shallow: the effect needs a self-assessing agent with a reference class, so any cross-domain invocation borrows the label for what is structurally plain anchoring, and it is the parent that is imported, not the effect. Crucially, when the cross-domain lesson — "rating something by its own absolute attributes, without consulting the reference class, produces distributionally incoherent comparison" — is genuinely wanted, it is already carried, in more general form, by anchoring_bias, because stripped of the self-rating paradigm and the incoherence diagnostic, the residue simply is that parent. And the effect's own interventions (reframe difficulty to flip the sign, supply a distribution to restore coherence) are self-judgment-specific and do not port. So the cross-domain reach belongs to anchoring_bias; the worse-than-average effect is the self-assessment instance that specializes it with a difficulty-indexed cross-over and a population-incoherence tell, and its distinctive cargo is exactly the part that does not travel. It clears the domain-specific bar comfortably within the self-assessment literature but sits below the prime bar — indeed close to home, given how thin its proprietary content is — because its only substrate-spanning content is already held by the prime it instantiates.

Relationships to Other Abstractions

Local relationship map for Worse-than-Average EffectParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Worse-than-AverageEffectDOMAINDomain-specific abstraction: Comparative Self-Assessment Crossover — is a kind ofComparative Sel…DOMAIN

Current abstraction Worse-than-Average Effect Domain-specific

Parents (1) — more general patterns this builds on

  • Worse-than-Average Effect is a kind of Comparative Self-Assessment Crossover Domain-specific

    The Worse-than-Average Effect is the below-average regime of the comparative self-assessment crossover, selected when a task is difficult, unfamiliar, or expert-dominated.

Hierarchy paths (5) — routes to 4 parentless roots

Not to Be Confused With

  • Better-than-average effect. The twin, not a separate phenomenon: on easy, familiar tasks most people rate themselves above average. It arises from the identical absolute-skill-anchoring mechanism; the two are one cross-over curve indexed by perceived difficulty, differing only in sign. Tell: is the task easy and near-universally practiced, placing most self-ratings above the estimated average (better-than-average), or hard and expert-dominated, placing them below (worse-than-average)? Same mechanism, opposite regime.

  • Dunning–Kruger effect. A skill-quartile claim: the least competent overestimate their ability most (and the most competent slightly underestimate), a relationship between actual skill and miscalibration. The worse-than-average effect is an average direction of self-rating set by task difficulty, holding across all competence levels. Tell: is the claim about who (by skill quartile) is most miscalibrated (Dunning–Kruger), or about which way the average rating tilts given the task's difficulty (worse-than-average)? Different independent variable.

  • Impostor syndrome. A felt, affective sense of fraudulence — competent individuals experiencing themselves as undeserving frauds despite evidence — often in high achievers. The worse-than-average effect is a comparative-rating direction produced by reference-class neglect, not an emotional self-experience. Tell: is it a persistent felt fear of being exposed as a fake (impostor syndrome), or a below-median self-placement on a hard task driven by absolute anchoring (worse-than-average)?

  • Regression to the mean. A statistical artifact of noisy repeated measurement — extreme scores drift toward the average on retest — with no self-assessment or anchoring involved. The worse-than-average effect is a systematic directional bias in a single comparative judgment. Tell: is the pattern a measurement/retest artifact of noisy data (regression), or a population's self-ratings systematically miscoded by skipping the reference class (worse-than-average)?

  • The parent prime anchoring_bias. The substrate-neutral core — a judgment anchored on the focal object's absolute attributes neglects the comparative distribution, yielding distributionally incoherent comparison — of which the worse-than-average effect and its better-than-average twin are the difficulty-modulated obverse halves. Outside a self-assessing agent, "worse-than-average" is just reference-class-neglect anchoring under a borrowed name. Tell: strip the self-rating paradigm and the population-incoherence tell and the residue simply is anchoring_bias (treated more fully elsewhere); the named effect is its self-assessment specialization.

Neighborhood in Abstraction Space

Worse-than-Average Effect sits in a crowded region of the domain-specific corpus (23rd percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.

Family — Social Perception & Self-Referential Bias (23 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12