Skip to content

Regression Discontinuity Design

Recover a causal effect from a threshold rule by comparing units just above and just below a sharp cutoff on a continuous running variable, where they are comparable in expectation, so any jump in the outcome at exactly the cutoff is attributable to the treatment rather than to selection.

Core Idea

Regression discontinuity design (RDD) is a quasi-experimental causal inference technique that exploits settings in which treatment is assigned by a sharp cutoff on a continuous running variable — a test score, an age threshold, an income line, a vote share — so that units just above the cutoff receive treatment and units just below do not. Because tiny fluctuations of the running variable near the cutoff are effectively uncontrolled by the subjects, units on either side of the threshold are comparable in expectation; any discontinuous jump in the outcome at exactly the cutoff is therefore attributable to the treatment, not to pre-existing differences. The design comes in two forms: sharp RDD, in which crossing the threshold deterministically assigns treatment (the probability of treatment jumps from 0 to 1 at the cutoff), and fuzzy RDD, in which crossing the threshold sharply increases the probability of treatment but does not determine it, in which case the cutoff functions as an instrumental variable and the estimand is a local average treatment effect for compliers. In both forms, the identifying assumption — that all determinants of the outcome other than treatment are smooth through the cutoff — is testable: covariate balance at the cutoff can be checked, and the McCrary density test detects whether subjects are bunching on one side of the threshold by manipulating their running variable, which would violate the as-if-random comparability of adjacent units. The founding application is Thistlethwaite and Campbell (1960) on the effect of National Merit Scholarship awards on later career outcomes; the design became a central tool of the credible-empirics revolution in econometrics and was part of the methodological contribution recognised by the 2021 Nobel Prize in Economic Sciences.

Structural Signature

Sig role-phrases:

  • the continuous running variable — a measured score (test score, age, income, vote share) along which units are ordered and near which small fluctuations are uncontrolled
  • the sharp cutoff — the precise administrative threshold on the running variable that assigns, or sharply shifts the probability of, treatment
  • the local as-if-randomization — the structural claim that units just above and just below the cutoff differ only by noise in the running variable, hence comparable in expectation
  • the smoothness assumption — the identifying requirement that every determinant of the outcome other than treatment passes continuously through the cutoff
  • the outcome discontinuity — the jump in the outcome at exactly the cutoff, attributed to treatment rather than selection
  • the falsification diagnostics — the McCrary density test for bunching/manipulation plus covariate-balance checks, by which credibility is earned rather than asserted
  • the sharp/fuzzy estimand classification — whether crossing deterministically assigns treatment (sharp: ATE at the cutoff) or only shifts its probability (fuzzy: cutoff as instrument, LATE for compliers)
  • the local-only scope (what it deliberately discards) — the effect is identified precisely at the cutoff and nowhere else, with bandwidth the explicit bias-precision knob; internal validity here does not extrapolate

What It Is Not

  • Not a randomised controlled trial. RDD finds as-if-random assignment in a threshold rule; it does not make it. Comparability holds only locally and only under an assumption — that every other determinant of the outcome is smooth through the cutoff — which an experimenter's randomisation would guarantee outright. The credibility is earned by diagnostics, not conferred by design.
  • Not a population-wide effect. The discontinuity identifies the treatment effect at the cutoff and nowhere else; units far from the threshold are not made comparable by the design. Internal validity here is precisely as narrow as it is clean, so a local estimate must not be read as the average effect across the whole population.
  • Not valid when subjects can manipulate the running variable. If borderline cases can precisely game their score, income, or other running variable to land on the favourable side, adjacent units are no longer as-if-randomly sorted and the outcome jump is confounded. This is a testable failure — the McCrary density test detects bunching — not an assumption to be waved through.
  • Not unconditionally credible. The identifying assumption can fail and is checkable: a jump in a pre-treatment covariate at the cutoff flags a confounded design. A clean-looking discontinuity is only a clean causal effect after covariate balance and density checks survive; the jump alone proves nothing.
  • Not always a deterministic-jump effect. When crossing the cutoff merely shifts the probability of treatment rather than determining it (fuzzy RDD), the threshold functions as an instrument and the estimand is a local effect for compliers, not the sharp average treatment effect. Reading a fuzzy design as if it delivered a deterministic jump misstates what was estimated.
  • Not "regression to the mean." The shared word "regression" is a coincidence of statistical vocabulary; RDD is a causal-identification design exploiting a threshold discontinuity, with no structural relation to the tendency of extreme measurements to move toward the average on remeasurement.

Scope of Application

Regression discontinuity design is a research method whose home is empirical causal inference on administratively-assigned treatments; its reach is across the substantive subfields that share that home, not across substrates — the design travels intact wherever a continuous running variable carries a sharp cutoff, only the score's referent changing. (The methodological lesson of found as-if-randomization rides on the natural-experiment parent; the threshold-discontinuity world-pattern recurs in control engineering and physics under their own frameworks, not as RDD travelling there.)

  • Education-policy evaluation — the founding turf: scholarship and school-entry cutoffs and grade-retention rules, the Thistlethwaite–Campbell National Merit case being the first application.
  • Health policy — sharp eligibility lines supply the discontinuity: the Medicare-at-65 threshold for coverage and mortality, BMI cutoffs for treatment eligibility.
  • Labour and welfare policy — benefit-eligibility ages and weeks-of-work thresholds, and means-tested income lines, assign treatment by crossing an administrative cutoff.
  • Crime and sentencing — mandatory-minimum thresholds on drug quantity or prior-conviction count create a jump in sentence at a precise cutoff.
  • Election studies — close-race vote-share discontinuities make which party takes office as-if-random in the narrow window around 50%.
  • Regulatory and corporate studies — firm-size, revenue, or emissions thresholds that trigger a regulation furnish the running variable and cutoff.
  • Development economics — poverty-line eligibility for cash transfers, and administrative- or geographic-boundary discontinuities, identify program effects where trials are infeasible.

Clarity

Naming RDD makes a class of causal opportunities legible that ordinary correlational analysis renders invisible: the sharp administrative thresholds scattered through policy — eligibility ages, score cutoffs, income lines, vote shares — stop being mere implementation details and become identification strategies hiding in plain sight. Where a pooled comparison of the treated and untreated is a confounded mess (scholarship winners differ from losers in ability and motivation; the over-65 differ from the under-65 in countless age-related ways), the design relocates the question to the immediate neighbourhood of the cutoff, where units differ only by uncontrollable small fluctuations of the running variable and are therefore comparable in expectation. The sharper question it licenses is local and answerable: at the threshold, is the outcome a smooth function of the running variable, or does it jump? — with any jump attributable to the treatment rather than to selection.

Its second clarity is that the design's central assumption is visible and testable rather than asserted. By stating the requirement crisply — every determinant of the outcome other than treatment must pass smoothly through the cutoff — RDD converts credibility into checkable diagnostics: covariate balance at the threshold, and the McCrary density test for whether subjects are bunching on one side by manipulating their running variable, which would betray that adjacent units are not as-if-randomly sorted. This also sharpens two distinctions practitioners must not blur. First, sharp versus fuzzy: when crossing the cutoff merely shifts the probability of treatment rather than determining it, the threshold becomes an instrument and the estimand is a local effect for compliers, not the deterministic jump. Second, local versus global: the design buys a clean effect precisely at the cutoff and nowhere else, so the concept keeps the analyst honest that internal validity here does not extend to units far from the threshold — the estimate is credible exactly where it is also narrow.

Manages Complexity

The general problem RDD confronts — recovering a causal effect from observational data where treatment was not randomised — is, taken whole, intractable: the treated and untreated differ on everything that drove their treatment status, and adjusting for that difference requires either knowing and measuring every confounder or defending an untestable claim that the ones left out do not matter. Scholarship winners differ from losers in ability and motivation; the over-65 differ from the under-65 in countless age-related ways; jurisdictions that vote one way differ from those that vote the other in everything that shaped the vote. RDD collapses this open-ended confounding problem to a single local question by exploiting one structural feature of the data-generating process — a sharp cutoff on a continuous running variable. Restricting attention to the immediate neighbourhood of the threshold, where units differ only by uncontrollable small fluctuations of the running variable and are therefore comparable in expectation, the entire tangle of confounders reduces to one checkable thing: at the cutoff, is the outcome a smooth function of the running variable, or does it jump? A jump is the treatment effect; the high-dimensional confounding problem has been traded for a one-dimensional question about behaviour at a point.

The compression is governed by a small, fixed set of parameters the analyst tracks rather than re-deriving causal warrant for each new policy setting. The identifying assumption itself reduces to a single requirement — every determinant of the outcome other than treatment passes smoothly through the cutoff — which in turn becomes a short, decidable diagnostic checklist: is there a running variable with a sharp cutoff; do pre-treatment covariates balance at the threshold; does the McCrary density test show no bunching that would betray manipulation of the running variable; is assignment at the cutoff sharp or fuzzy. From those, both the validity and the meaning of the estimate read off along a clean branch structure. If covariates balance and density is smooth, the as-if-random comparability holds and the jump is a credible causal effect; if covariates jump or subjects bunch on one side, the design fails and the discontinuity is confounded. Whether the cutoff deterministically assigns treatment (sharp) or merely shifts its probability (fuzzy) sets the estimand: a deterministic jump in the first case, a local effect for compliers — with the cutoff serving as an instrument — in the second. And the scope is fixed by the same structure: the effect is credible at the cutoff and nowhere else, so the analyst reads internal validity and its narrow reach off the design simultaneously, never mistaking a clean local estimate for a global one. Instead of confronting the full causal-inference problem afresh in education, health, labour, or elections, the analyst locates the threshold, runs the balance and density checks, classifies sharp versus fuzzy, and reads off both whether the effect is identified and exactly which units it is identified for — an intractable confounding problem reduced to a local jump plus a short checklist with a decidable branch structure.

Abstract Reasoning

The founding move is spotting found randomization in a threshold rule — reasoning from the structure of an administrative system to a local natural experiment hiding in plain sight. The analyst scans a policy setting for a sharp cutoff on a continuous running variable — a test score, an eligibility age, an income line, a vote share — and infers that near that cutoff, where the running variable's small fluctuations are effectively uncontrolled by subjects, units on either side are comparable in expectation. The characteristic inference runs from "treatment is assigned by crossing this threshold" to "adjacent units differ only by noise in the running variable, so any jump in the outcome at the cutoff is the treatment effect." The move trades an intractable global confounding problem (winners differ from losers on everything that drove their treatment) for a single local question — at the cutoff, is the outcome smooth, or does it jump? — answerable where pooled comparison is hopeless.

The decisive move is converting the identifying assumption into testable diagnostics and running them. Rather than assert as-if-randomness, the analyst states the requirement crisply — every determinant of the outcome other than treatment must pass smoothly through the cutoff — and reasons to two falsification checks. A manipulation diagnostic: if subjects can precisely game their running variable to land just above or below (nudging borderline scores, gaming reported income), the comparability fails, and the McCrary density test detects this by looking for bunching on one side of the threshold. A balance diagnostic: pre-treatment covariates should be continuous through the cutoff, so a jump in a covariate flags a confounded design. The inference runs from an observed bunching or covariate discontinuity to the as-if-random comparability is broken, so the outcome jump is not a clean effect — and from smooth density plus balanced covariates to the design holds. The move's discipline is that credibility is earned by surviving checks, not claimed.

A classification move sets the estimand by reading the assignment mechanism at the cutoff. The analyst asks whether crossing the threshold deterministically assigns treatment (probability jumps 0 to 1) or merely shifts its probability, and infers the meaning of the jump accordingly: sharp RDD recovers the average treatment effect at the cutoff, while fuzzy RDD makes the cutoff an instrument and recovers a local effect for compliers — those whose treatment status was actually changed by crossing. The inference runs from the shape of the treatment-probability step to what the discontinuity estimates, and the failure mode it prevents is reading a fuzzy-design jump as if it were a deterministic effect.

Finally, a scope-bounding move keeps internal and external validity apart at the point of estimation. The analyst reasons that the design buys a clean effect precisely at the cutoff and nowhere else — units far from the threshold are not made comparable by it — so the estimate is credible exactly where it is also narrow. The inference runs from the comparability holding only in the immediate neighbourhood to this effect is local and may not extrapolate, with the bandwidth choice (how wide a window still counts as "local") an explicit knob trading bias against precision. The move forbids mistaking a clean local estimate for a population-wide one, so the analyst reports the effect together with the narrow population it is identified for.

Knowledge Transfer

As with any identification strategy, what transfers is a research design — a procedure for extracting a causal effect from data the analyst did not randomise — not a structural pattern in the world. With that fixed, RDD's within-domain transfer is wide and literal, because its precondition is exact and substrate-neutral: a continuous running variable with a sharp administrative cutoff that assigns (or sharply shifts the probability of) treatment. Wherever that shape recurs across empirical social science the same design is run with the same checklist — locate the cutoff, define the running variable, restrict to a local window, run the McCrary density test for manipulation, check covariate balance, classify sharp versus fuzzy, estimate the jump. It is the founding tool of education-policy evaluation (scholarship and school-entry cutoffs, the Thistlethwaite–Campbell case), and it operates identically in health policy (the Medicare-at-65 threshold, BMI eligibility lines), labour and welfare (benefit-eligibility ages and weeks-of-work thresholds, means-tested income lines), crime and sentencing (mandatory-minimum quantity and prior-conviction thresholds), election studies (close-race vote-share discontinuities), regulatory and corporate work (firm-size and emissions thresholds), and development economics (poverty-line cash-transfer eligibility, administrative-boundary discontinuities). The transfer is literal rather than analogical because it is one design solving one problem-shape: the vocabulary travels untranslated (running variable, cutoff, local effect, bunching, sharp/fuzzy, bandwidth), only the referent of the score changing, so an analyst fluent in scholarship-cutoff RDD reads a Medicare or vote-share RDD with no relearning.

The honest boundary is that this wide reach is spread across substantive fields within one methodological home — empirical causal inference on administratively assigned treatments — and the design does not export to natural science, engineering, mathematics, or the humanities as such; where it appears there at all, it is imported as an analytic tool, not rediscovered. When a cross-domain lesson is genuinely wanted, two things carry, neither of them "RDD" by name. The first is the shared abstract mechanism it instantiates as a method: found, as-if-random assignment exploited without an experimenter — the natural-experiment / quasi-experimental identification family (with fuzzy RDD literally a special case of instrumental variables), which is the level at which the credible-empirics move is substrate-portable. The second is the more general world-pattern that a sharp threshold rule creates a local discontinuity that can be diagnosed — a structural property of threshold-based systems that genuinely recurs as hysteresis-around-a-threshold in control engineering and as phase-transition diagnostics in physics, but those domains have their own home-grown frameworks and the recurrence is co-instantiation of the general threshold-discontinuity pattern, not RDD travelling to them. The home-bound cargo is everything that makes RDD specifically itself: the running-variable-and-cutoff identification logic, the McCrary manipulation test, the sharp-versus-fuzzy estimand distinction, the complier-LATE interpretation, the bandwidth/bias trade, and the local-only validity bookkeeping — none of which has a referent where there is no treatment to identify. So calling any threshold effect in a physical system "a regression discontinuity" is analogy: it borrows the jump-at-a-cutoff shape while dropping the causal-identification machinery that is the whole point. The disciplined position is that the design transfers across the policy-evaluation cluster as genuine shared machinery, the methodological lesson rides on the natural-experiment parent, the world-pattern rides on the general threshold-discontinuity structure, and the named RDD apparatus stays in its causal-inference home (see Structural Core vs. Domain Accent).

Examples

Canonical

The founding application is Thistlethwaite and Campbell (1960). They studied whether receiving public recognition — a Certificate of Merit in the National Merit Scholarship program — affected students' later attitudes and career aspirations. Winners plainly differed from non-winners in ability, so a naive comparison was hopelessly confounded. But the award was assigned by a sharp cutoff on a qualifying test score. Thistlethwaite and Campbell reasoned that students scoring just above and just below that cutoff were essentially indistinguishable in ability — the difference of a point or two being noise — yet only those above received the certificate. Plotting later outcomes against the test score, they looked for a discontinuous jump at exactly the cutoff. A jump there could be credited to the award itself, not to the ability differences a pooled winner-versus-loser contrast would confound.

Mapped back: The qualifying test score is the continuous running variable; the award threshold is the sharp cutoff. The claim that students a point or two apart are indistinguishable is the local as-if-randomization, resting on the smoothness assumption that ability and everything else pass continuously through the cutoff. A jump in later outcomes at exactly that score is the outcome discontinuity, credited to the award rather than to selection.

Applied / In Practice

A widely cited real deployment is Card, Dobkin, and Maestas's study of Medicare eligibility, which begins abruptly at age 65 in the United States. Because turning 65 is not something people can nudge, and because those aged 64 years 11 months and 65 years 1 month are otherwise nearly identical, the age-65 threshold supplies as-if-random variation in insurance coverage. The researchers plotted health-insurance coverage and healthcare utilization against age and found sharp jumps at exactly 65: coverage rose and use of hospital and physician services increased discontinuously. Because little else about people changes discontinuously at that birthday, the jumps identify the causal effect of gaining Medicare. The estimate is, by construction, local — an effect for people near 65 — and the design's credibility rests on the smoothness of everything else through the cutoff, which is checkable in the data.

Mapped back: Age is the continuous running variable and 65 is the sharp cutoff; that a person cannot precisely game their birthday underwrites the local as-if-randomization. The abrupt rise in coverage and utilization at 65 is the outcome discontinuity. That the result speaks only to people near 65 is the local-only scope, and its trustworthiness turns on the smoothness assumption being borne out by the covariate and density diagnostics.

Structural Tensions

T1: Clean identification versus narrow reach (credible exactly where it is least generalizable). RDD's signature virtue is that comparability holds precisely at the cutoff, so the jump is a clean causal effect there — and its signature limitation is that this holds only there. The effect is credible exactly where it is also narrow; the same locality that severs confounding forbids extrapolation to units far from the threshold. A Medicare-at-65 estimate is trustworthy for people near 65 and silent about the effect of coverage at 45. The design buys internal validity at the direct cost of external reach, and there is no free widening: pushing the bandwidth outward to say more about the population reintroduces the very heterogeneity the locality excluded. Diagnostic: Is the policy question actually about units at the threshold, or is a clean local estimate being stretched to a population it was never identified for?

T2: Bandwidth as bias-precision knob (the window that adds data also adds bias). "Local" is not self-defining; the analyst must choose how wide a window still counts as near the cutoff, and that bandwidth is an explicit knob. Narrow it and the comparability is purer but the sample shrinks toward noise; widen it and precision improves but units increasingly differ on more than running-variable fluctuation, so bias creeps in. The estimate genuinely moves with the choice, which means an analyst has a degree of freedom that can be exercised — innocently or not — to produce a preferred discontinuity. The design's credibility rests on a parameter it cannot itself fix, and honesty requires showing the estimate is stable across reasonable bandwidths rather than reporting the one that looks best. Diagnostic: Does the discontinuity survive across a defensible range of bandwidths, or does it appear only at a window the analyst had reason to prefer?

T3: Testable assumptions versus untestable smoothness (diagnostics that survive still do not prove the design). RDD's advance over asserted causal claims is that its identifying assumption is checkable — McCrary density for manipulation, covariate balance at the cutoff. But passing those checks is necessary, not sufficient: the core assumption is that every determinant of the outcome other than treatment is smooth through the cutoff, and an unobserved determinant that jumps at exactly the threshold leaves no footprint in the covariates you happened to measure or in the density. The diagnostics can all pass while a coincident policy — another program that also switches on at 65, another rule that also triggers at the same score — rides the same cutoff. The visible testability of the assumption can breed a false security that the assumption is confirmed rather than merely not-yet-refuted. Diagnostic: Is there any other treatment, rule, or determinant that also changes at exactly this cutoff, which the balance and density tests would not detect?

T4: Sharp precision versus fuzzy prevalence (the cleaner form is the rarer one). Sharp RDD, where crossing deterministically flips treatment 0 to 1, yields the cleanest interpretation — the ATE at the cutoff. But truly deterministic cutoffs are relatively rare; most real thresholds only shift the probability of treatment, which is fuzzy RDD, where the cutoff becomes an instrument and the estimand contracts to a local effect for compliers. So the form that is easiest to interpret is the form least often available, and the pressure to treat an approximately-sharp design as fully sharp — ignoring the non-compliers who got or missed treatment despite their side of the line — misstates what was estimated. The interpretive clarity of the sharp case is a temptation to overclaim it where reality is fuzzy. Diagnostic: Does crossing this cutoff determine treatment, or only raise its probability — and is the reported estimand honestly the complier-LATE rather than a population ATE?

T5: Found-randomization opportunism versus manipulable thresholds (the cutoff that assigns treatment invites gaming). The design's brilliance is spotting as-if-random assignment hiding in an administrative rule. But the moment a threshold carries a real consequence, the agents subject to it acquire a motive to sort across it — nudging a reported income under an eligibility line, gaming a borderline test score, timing an action to land on the favorable side. The very salience that makes a cutoff worth studying is what gives subjects reason to manipulate the running variable, destroying the comparability of adjacent units. The McCrary test catches gross bunching, but subtle, partial sorting near a high-stakes threshold can degrade the design below detection. The consequential cutoffs most worth exploiting are the ones most exposed to the manipulation that invalidates them. Diagnostic: Do the units near this cutoff have both the motive and the fine-grained control to sort across it, and would the resulting sorting be coarse enough for the density test to see?

T6: Autonomy versus reduction (its own named design or the causal-inference instance of its parents). RDD is a named, canonically studied design with proprietary apparatus — the running-variable/cutoff identification logic, the McCrary manipulation test, the sharp-versus-fuzzy estimand distinction, the complier-LATE interpretation, the bandwidth trade — and that apparatus transfers literally across the policy-evaluation cluster (education, health, labour, elections, development), where only the score's referent changes and the vocabulary travels untranslated. But it does not export to physics or engineering as mechanism; where a cross-domain lesson is wanted, two more general things carry — the natural-experiment / quasi-experimental identification family (with fuzzy RDD literally a special case of instrumental variables), which is where the credible-empirics move is substrate-portable, and the world-pattern that a sharp threshold rule creates a diagnosable local discontinuity (recurring as hysteresis in control engineering, phase-transition diagnostics in physics) under those fields' own frameworks. Calling a physical threshold effect "a regression discontinuity" borrows the jump-at-a-cutoff shape while dropping the causal-identification machinery that is the whole point. The tension is between a standalone design that earns its own apparatus and the recognition that its portable cargo belongs to the natural-experiment parent and the general threshold-discontinuity pattern. Diagnostic: Resolve toward the natural-experiment / IV parent (and the threshold-discontinuity world-pattern) when asking what travels beyond policy evaluation; toward "RDD" when running the balance, density, and sharp/fuzzy checklist on an actual administrative cutoff.

Structural–Framed Character

Regression discontinuity design lands at the mixed midpoint of the structural–framed spectrum — a neutral research design that transfers literally within its methodological home, but more domain-bound than a broad instrument like regression and so short of mixed-structural. On evaluative_weight it patterns structural: "RDD" convicts and praises nothing — it names an identification strategy, a way to extract a causal effect, and the whole analysis is evaluatively neutral. On import_vs_recognize it patterns structural within its cluster: the entry insists the within-domain transfer is literal — "one design solving one problem-shape," the vocabulary (running variable, cutoff, bunching, sharp/fuzzy, bandwidth) travelling untranslated across education, health, labour, elections, and development, "only the referent of the score changing." That literal re-use across substantive fields is recognition of the same machinery, not import-by-analogy, and it is what gives the entry real structural ballast.

What pulls it toward the framed side, and keeps it a notch more framed than regression, is threefold. On human_practice_bound it patterns framed in the instrument sense: RDD does not run in nature observer-free — there is no "regression discontinuity design" in the world without an analyst who spots the found randomization, restricts to a local window, and runs the diagnostics; it is a research procedure, a thing done. (The world-pattern that genuinely recurs in nature — a sharp threshold producing a local discontinuity — is the general threshold-discontinuity structure, not RDD, which is precisely the entry's point.) On institutional_origin it is a framed artifact of a tradition: the named design, the McCrary manipulation test, the sharp-versus-fuzzy estimand distinction, the complier-LATE interpretation, the bandwidth/bias bookkeeping, and the whole credible-empirics apparatus (Thistlethwaite–Campbell 1960, the 2021 Nobel) are furniture of econometrics methodology. On vocab_travels it is mixed-leaning-framed and more pinned than regression: where regression's precondition is broad enough that its vocabulary reads over any data with a modelled systematic part, RDD's precondition is narrow — a continuous running variable carrying a sharp administrative cutoff that assigns treatment — so its vocabulary travels literally only across the empirical-causal-inference/policy-evaluation cluster and does not export as a design to physics or engineering; where a threshold effect appears there, calling it "a regression discontinuity" is analogy that drops the causal-identification machinery.

The portable structural content genuinely resolves into two levels, which is why more than one skeleton is warranted here. The primary methodological skeleton is found, as-if-random assignment exploited without an experimenter — the natural-experiment / quasi-experimental identification family, with fuzzy RDD literally a special case of instrumental_variables; that is the level at which the credible-empirics move is substrate-portable, and it is what RDD instantiates from its parent, not what makes "RDD" itself travel. The secondary skeleton is the world-pattern that a sharp threshold rule creates a diagnosable local discontinuity, which co-instantiates in control-engineering hysteresis and physical phase-transition diagnostics under those fields' own frameworks — a general structure RDD shares, not RDD reaching there. The cross-domain reach belongs to the natural-experiment/IV parent and the threshold-discontinuity pattern; RDD keeps for itself the home-bound cargo — the running-variable/cutoff logic, the manipulation and balance tests, the sharp/fuzzy distinction, the local-only validity bookkeeping — that has no referent where there is no treatment to identify. Its character: an evaluatively neutral causal-inference design that transfers literally across the policy-evaluation cluster, structural in the found-as-if-randomization skeleton it instantiates from its natural-experiment/IV parent, but a specialized econometrics instrument whose apparatus and narrow precondition keep it mixed rather than mixed-structural.

Structural Core vs. Domain Accent

This section decides why regression discontinuity design is a domain-specific abstraction and not a prime, and it carries the case for its domain-specificity — no other section does. Like other instruments, RDD transfers by literal re-use rather than metaphor within its home, so the not-a-prime verdict turns on the causal-identification apparatus bundled into the named design and on its narrow precondition.

What is skeletal (could lift toward a cross-domain prime). Strip the econometrics away and the portable content resolves into two distinct levels, which is why more than one skeleton is warranted. The primary methodological skeleton is found, as-if-random assignment exploited without an experimenter — a natural experiment recognized in the structure of an administrative rule and used to identify a causal effect the analyst did not randomize; fuzzy RDD is literally a special case of using the cutoff as an instrument. The secondary skeleton is a world-pattern: a sharp threshold rule creates a local discontinuity that can be diagnosed at the cutoff. The first is genuinely substrate-portable at the level of the natural-experiment / quasi-experimental identification family (with instrumental_variables as the fuzzy special case); the second co-instantiates in control-engineering hysteresis and physical phase-transition diagnostics under those fields' own frameworks. Both are the cores it shares, not what makes RDD distinctive — and the second is a pattern RDD shares, not RDD reaching into physics.

What is domain-bound. Everything that makes it regression discontinuity design in particular is causal-inference furniture and none of it survives extraction. It presupposes a treatment to identify assigned by a continuous running variable crossing a sharp administrative cutoff. On that sit the worked apparatus: the running-variable/cutoff identification logic; the smoothness assumption that every other determinant of the outcome passes continuously through the cutoff; the falsification diagnostics (the McCrary density test for manipulation/bunching, covariate-balance checks); the sharp-versus-fuzzy estimand distinction and the complier-LATE interpretation; the bandwidth bias-precision knob; and the local-only validity bookkeeping. These are the empirical cases and craft vocabulary of quasi-experimental econometrics (Thistlethwaite–Campbell 1960, the credible-empirics revolution, the 2021 Nobel). The decisive test: remove the treatment-to-identify — as in a control system's hysteresis or a physical phase transition, where a threshold produces a jump but nothing is being estimated — and the identification machinery has no referent; calling that a "regression discontinuity" borrows the jump-at-a-cutoff shape while dropping the whole point. None of the apparatus means anything where there is no causal effect to recover.

Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. RDD's transfer is bimodal. Within its methodological home — empirical causal inference on administratively assigned treatments — it travels literally across education, health, labour, crime, elections, regulation, and development, because each supplies the one thing it needs, a continuous running variable with a sharp cutoff, so the checklist (locate cutoff, restrict to a window, run McCrary and balance, classify sharp/fuzzy, estimate the jump) moves untranslated, only the score's referent changing. But that literal reach is spread across substantive fields inside one home, not across substrates. Beyond that home it travels only by analogy: a threshold effect in a physical system borrows the shape without the identification logic. And when the bare structural lesson is genuinely wanted cross-domain, it is already carried, in more general form, by the parent RDD instantiates — the natural-experiment / quasi-experimental identification family (with fuzzy RDD a special case of instrumental_variables) for the methodological move, and the general threshold-discontinuity world-pattern for the physical recurrence. The cross-domain reach belongs to those; RDD, as named, keeps the running-variable/cutoff logic, the manipulation and balance tests, the sharp/fuzzy distinction, the bandwidth trade, and the local-only bookkeeping that should stay home. (Note too the pure name-collision with regression to the mean, which shares only the word.)

Relationships to Other Abstractions

Local relationship map for Regression Discontinuity DesignParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.RegressionDiscontinuity DesignDOMAINDomain-specific abstraction: Natural Experiment — is a kind ofNaturalExperimentDOMAIN

Current abstraction Regression Discontinuity Design Domain-specific

Parents (1) — more general patterns this builds on

  • Regression Discontinuity Design is a kind of Natural Experiment Domain-specific

    Regression discontinuity design is a natural experiment specialized to found as-if-random assignment at a sharp cutoff on a continuous running variable.

Hierarchy paths (12) — routes to 8 parentless roots

Not to Be Confused With

  • Regression (the statistical method). The general fitting procedure — model an outcome as a function of predictors under a loss. RDD borrows regression's "regression" and uses it to estimate the jump, but RDD is an identification design: its content is the found as-if-randomization at a cutoff, not the curve-fitting. A regression is run in service of it, but any well-fit regression is not an RDD. Tell: is the point to fit a systematic function of predictors, or to exploit a sharp threshold so that a jump at the cutoff is attributable to treatment?
  • Instrumental variables (IV). The broader identification strategy that uses an exogenous variable affecting treatment but not the outcome directly. Fuzzy RDD is literally a special case — the cutoff serving as the instrument, recovering a complier-LATE. Part-vs-whole: IV is the general family, fuzzy RDD is the threshold-based instance; sharp RDD dispenses with the instrument entirely. Tell: is the exogenous variation a general instrument, or specifically a probability-of-treatment jump at a running-variable cutoff?
  • Difference-in-differences. A sibling quasi-experimental design that identifies effects from differential change over time between a treated and control group under a parallel-trends assumption. RDD identifies from cross-sectional proximity to a cutoff under a smoothness assumption; no time dimension or parallel trends is required. Tell: does identification rest on comparing trends before/after across groups, or on comparing units just above/below a threshold?
  • Interrupted time series. The close neighbour where the discontinuity is in time — an outcome series jumps when an intervention switches on at a date. RDD's discontinuity is in a continuous running variable (a score, age, income) across units, not a break in a temporal sequence. This is the easiest confusion because both are "discontinuity designs." Tell: is the running variable time (a before/after break), or a cross-sectional score with units sorted on either side of a cutoff?
  • Matching / propensity-score methods. A sibling observational-causal approach that balances treated and control units on measured covariates across the whole sample. RDD instead achieves local balance by design through as-if-random sorting near the cutoff, and makes no whole-sample comparability claim — its estimate is credible only at the threshold. Tell: is comparability engineered by conditioning on observed covariates over the population, or by restricting to a local window where units differ only by running-variable noise?
  • The natural-experiment / quasi-experimental identification family it instantiates (with instrumental_variables as the fuzzy special case). The substrate-portable parent — found, as-if-random assignment exploited without an experimenter — plus the general world-pattern that a sharp threshold creates a diagnosable local discontinuity (recurring as hysteresis in control engineering, phase-transition diagnostics in physics, under their own frameworks). Those, not "RDD," carry the cross-domain lesson; calling a physical threshold effect a "regression discontinuity" is analogy that drops the causal-identification machinery. Tell: is there a treatment to identify? No treatment, no RDD — only the bare threshold-discontinuity pattern, treated fully in a later section.

Neighborhood in Abstraction Space

Regression Discontinuity Design sits in a sparse region of the domain-specific corpus (89th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (309 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12