Skip to content

Riskiest Assumption Test

Rank a plan's assumptions by consequence-if-false times uncertainty times upstream-position, then spend the next effort on the cheapest credible test of the top one — the question whose answer would most reduce total wasted work.

Core Idea

The riskiest assumption test (RAT), popularised in lean-startup methodology and structurally related to the "killer experiment" in pharmaceutical R&D and "critical-path test" in engineering, is the decision discipline of identifying — from the full set of assumptions a plan depends on — the single assumption that combines (a) the largest downside if false with (b) the team's weakest independent evidence for it, and then designing the cheapest credible experiment to test that assumption before committing downstream resources. The structural commitment is to spend the next unit of effort on the question whose answer would most reduce total expected wasted effort across the plan's dependency graph.

The underlying calculation is a product of three terms: the consequence of the assumption being false (how much downstream work is invalidated), the probability of being wrong (how weak the current evidence is), and the upstream position (how many subsequent commitments depend on the assumption being true). An assumption can be highly consequential but well-supported, making it a low priority; can be poorly supported but mild in failure, equally a low priority; or can be both consequential and poorly supported but downstream in the dependency graph, where it can be deferred until upstream tests are resolved. The discipline is a priority ordering, not merely "test early": it specifically targets the assumption at the top of the consequence-times-uncertainty-times-upstream-position ranking, even when that assumption is socially inconvenient to test, requires an experiment outside the team's comfort zone, or reveals an uncomfortable answer. The convenience trap — running the test the team knows how to run rather than the test whose answer would most change the plan — is the named failure mode the RAT discipline is designed to prevent.

Structural Signature

Sig role-phrases:

  • the dependency-graphed plan — a coupled set of intended actions whose downstream commitments rest on upstream assumptions
  • the assumption inventory — the full set of propositions the plan would fail without, surfaced for scoring
  • the consequence term — how much downstream work each assumption invalidates if it is false
  • the uncertainty term — how weak the team's current independent evidence for each assumption is (the unknown-distribution endpoint, not known exposure)
  • the upstream-position term — how many later commitments depend on the assumption holding
  • the consequence × uncertainty × upstream ranking — the product that orders the assumptions and names the single top entry
  • the cheapest credible falsification — the minimum-cost experiment that could disconfirm the top-ranked assumption, run before downstream resources are sunk
  • the ranking-over-tooling discipline — the commitment that the ranking, not the team's existing instruments, governs sequence, so a test's awkwardness is irrelevant to its priority
  • the convenience trap — the named failure mode: running the test the team knows how to run rather than the one whose answer would most change the plan

What It Is Not

  • Not "test early" or "de-risk everything." RAT is a priority ordering, not a blanket injunction to validate. Its content is the ranking — consequence × uncertainty × upstream-position — that picks the single assumption to test next; an undifferentiated "test early" gives no sequence and collapses in practice to testing whatever is nearest to hand, which is precisely the failure RAT names.
  • Not "test the most consequential assumption." A consequential assumption that is already well-supported ranks low, because there is little uncertainty to resolve. The target is the assumption that is both high-consequence and poorly evidenced and upstream — high stakes alone do not promote it, and weak evidence alone does not either.
  • Not running the test the team knows how to run. The defining discipline is that the ranking, not the team's tooling, governs sequence — testing for evidence, not testing for comfort. A test's awkwardness, unfamiliarity, or threat of an uncomfortable answer is irrelevant to its priority; drifting toward the convenient test is the convenience trap the discipline exists to defeat.
  • Not an MVP. A minimum viable product is one instrument for running the test in product development, not the ordering discipline itself. RAT is about which assumption to test; the MVP is one delivery vehicle, and "MVP everything" loses the ranking that is RAT's whole content.
  • Not the formal value-of-information calculation. RAT informally instantiates value-of-information, but as practiced it is a ranking heuristic on three estimable terms, not a computation over explicit utilities and priors. It is also not critical-path analysis: critical path scores time slippage in a schedule, while RAT scores truth slippage — which false belief most invalidates the plan — even though the two often coincide.

Scope of Application

The riskiest assumption test names a planning discipline whose home is lean-startup product development, but the ordering it codifies — invest the next unit of effort where its result would most change downstream decisions — recurs as the same mechanism, under its own name, across several planning fields; the map below is those genuine co-instances. The boundary is that the named "RAT" cargo (the MVP instruments, the startup codification) stays home-bound, and the cross-field lesson is carried by the parent (value_of_information), so loose borrowings flattened to "do early tests" or "MVP everything" are over-reading, not habitat.

  • Lean-startup product development — the named home: design the cheapest experiment (a presales page, a concierge test) that could falsify the riskiest product assumption before scaling, and pivot if it fails.
  • Pharmaceutical R&D ("killer experiment" / kill-fast pipeline) — early, cheap, hard-to-pass tests on the single assumption most likely to fail (efficacy in the predictive model, toxicity profile), run precisely to avoid the far larger downstream cost of a Phase III failure.
  • Military and operational planning (red-teaming, pre-mortem) — rehearsals and intelligence collection that test the riskiest assumption — usually about adversary response, terrain, or coalition reliability — before committing the maneuver.
  • Engineering and aerospace (critical-path testing) — under the technology-readiness-level framework, targeting the assumption whose failure would invalidate the design before integration, rather than the assumption easiest to test.
  • Experimental science (Platt's strong inference) — designing the experiment that maximally discriminates between consequential competing hypotheses, not the one most likely to "work."
  • Investment and credit due diligence — allocating the diligence budget to the single assumption that, if false, would most reverse the thesis (regulatory approval, founder retention, channel economics), not the most easily investigated one.
  • Disaster preparedness (tabletop exercises) — testing the riskiest unverified assumption in the response plan (mutual-aid availability, communications resilience) rather than rehearsing the most familiar scenario.

Clarity

Naming the riskiest assumption test makes legible a distinction that planning discussions otherwise hide: testing for evidence versus testing for comfort. Without the concept, "we should validate early" sounds like sound discipline, yet the tests a team actually runs drift toward the ones it is already set up to run — the engineers build, because building is what they know — and the gap between "the test we can run" and "the test whose answer would most change the plan" stays invisible. RAT names that gap as the convenience trap and forces the sharper question: not are we testing? but of everything this plan depends on, which single assumption sits at the top of the consequence × uncertainty × upstream-position ranking, and is that the one we are about to test? That reframes the unit of planning effort from an activity to be scheduled into a scarce resource to be aimed at the question whose answer most reduces total expected wasted work.

It also sharpens a second distinction the field routinely blurs: risk — exposure under a known distribution of outcomes — versus uncertainty, where the distribution itself is unknown. RAT targets the uncertainty endpoint specifically: the priority is not the largest exposure under known statistics but the assumption that is least well-characterized and most consequential, which is precisely the cell ordinary risk-management vocabulary cannot see. And the three-term ranking does discriminating work that "test early" cannot, by separating cases that look alike from the outside: a consequential but well-supported assumption (low priority — little uncertainty to resolve), a poorly-supported but mild-failure assumption (low priority — little consequence), and a consequential, poorly-supported, but downstream assumption (deferrable until its upstream dependencies resolve). Making that ordering explicit renders the prioritization auditable, where intuitive planning leaves it tacit — and it tells the team that the awkwardness of a test is no reason to deprioritize it, because declining to run the top-ranked test now is a decision to discover its failure later, under worse conditions.

Manages Complexity

A plan of any size rests on a long, unranked list of assumptions — the team can build it, customers will pay the price, the channel will convert, acquisition cost will hold, the partner will deliver, the regulation will permit — and the naive response is an overwhelming injunction to "de-risk everything," which gives no guidance on sequence and collapses in practice to testing whatever is nearest to hand. RAT compresses that whole space to a single ordering question by scoring every assumption on the same three terms — consequence if false (how much downstream work it invalidates), uncertainty (how weak the current evidence is), and upstream position (how many later commitments depend on it) — and directing the next unit of effort at whatever sits atop the product of the three. The high-dimensional "what could go wrong, and in what order do we check?" thereby reduces to tracking one ranking and reading the top entry off it. The three terms also supply a clean branch structure that resolves cases an undifferentiated "test early" cannot tell apart: a consequential but well-supported assumption ranks low because there is little uncertainty to resolve; a poorly-supported but mild-failure assumption ranks low because there is little consequence; a consequential and poorly-supported but downstream assumption is deferred until its upstream dependencies resolve. Only the cell that is simultaneously high-consequence, high-uncertainty, and upstream rises to the top, and the analyst reads the next experiment to design directly off that cell rather than re-deriving the plan's whole risk profile. The compression's signature discipline is that the ranking, not the team's tooling, governs sequence: by making consequence × uncertainty × upstream-position explicit it renders the prioritization auditable and defeats the convenience trap — running the test the team knows how to run rather than the one whose answer would most change the plan — so the awkwardness of a test drops out of the ordering entirely. What was an unmanageable cloud of contingencies becomes a one-dimensional priority read off a single product of three estimable terms.

Abstract Reasoning

The riskiest assumption test licenses a set of reasoning moves built on one structural calculation — score each assumption by consequence × uncertainty × upstream-position and aim the next unit of effort at the top entry — and on the discipline that the ranking, not the team's tooling, governs sequence.

Diagnostic — rank a plan's assumptions on three terms and read the next experiment off the top cell, distinguishing cases an undifferentiated "test early" cannot tell apart. The characteristic move scores every proposition the plan would fail without on the same three estimable terms: consequence if false (how much downstream work it invalidates), uncertainty (how weak the current evidence is), and upstream position (how many later commitments depend on it). The analyst then reads the next experiment to design directly off the highest-scoring cell rather than re-deriving the plan's whole risk profile. The three terms do discriminating work: a consequential but well-supported assumption ranks low (little uncertainty to resolve); a poorly-supported but mild-failure assumption ranks low (little consequence); a consequential and poorly-supported but downstream assumption is deferred until its upstream dependencies resolve. Only the cell that is simultaneously high-consequence, high-uncertainty, and upstream rises to the top — so the inference runs from "here is the unranked cloud of what could go wrong" to "here is the single assumption whose answer would most reduce total expected wasted effort, and therefore the test to run next."

Diagnostic of the failure mode — detect and name the convenience trap by comparing the test the team can run against the test the ranking demands. The concept's sharpest move is to surface a gap planning discussions hide: the difference between testing for evidence and testing for comfort. The analyst watches where a team's tests actually drift — toward what it is already set up to run (the engineers build, because building is what they know) — and diagnoses the convenience trap whenever the test being run is not the test atop the consequence × uncertainty × upstream ranking. The move is to treat the awkwardness or unfamiliarity of a test as irrelevant to its priority: a top-ranked assumption that is socially inconvenient, requires an experiment outside the team's comfort zone, or threatens an uncomfortable answer is still the one to test, and a plan that defers it because it is awkward is diagnosed as having substituted feasibility for expected harm-of-being-wrong.

Interventionist — run the cheapest credible test of the top-ranked assumption now, and predict the cost of deferral. The corrective is to design the minimum-cost experiment that could falsify the top-ranked assumption and run it before committing downstream resources — a presales campaign with mock pricing before building, an efficacy test in the predictive model before scale-up, a rehearsal of the assumed adversary response before committing the maneuver. The kill-fast prediction is structural, not merely prescriptive: declining to run the top-ranked test now is a decision to discover its failure later, under worse conditions, after downstream commitments that depended on it have already been sunk. So the move forecasts that effort spent on lower-ranked or convenient tests while the top assumption stays unexamined is effort at risk of being wholly invalidated, and that the cheapest credible falsification-attempt, run first, dominates any later testing schedule.

Boundary-drawing — target the uncertainty endpoint, separate truth-slippage from time-slippage, and distinguish the meta-procedure from the instruments and statistics it deploys. Several lines the concept draws. First, risk versus uncertainty: the priority is not the largest exposure under a known distribution but the assumption least well-characterized and most consequential — the cell ordinary risk-management vocabulary cannot see — so the move is to aim at the unknown-distribution endpoint, not the well-quantified-exposure one. Second, the boundary against schedule reasoning: the ranking is about truth slippage in a set of assumptions (which false belief most invalidates the plan), distinct from time slippage in a schedule (which delayed task most slips the deadline), though the two often coincide — so the move is to score by consequence-of-being-wrong, not by duration. Third, the meta-procedure boundary: the riskiest assumption test concerns which experiment to run, distinct from the statistical procedure run inside any one experiment, and distinct from its delivery instruments (a minimum viable product is one way to run the test, not the ordering discipline itself) and its input techniques (imagining failure supplies candidate assumptions, but the testing-priority commitment is the added move). So the analyst scopes the analysis to the ranking-and-test discipline and does not conflate it with the statistics, the instrument, or the failure-imagining that feed it.

Knowledge Transfer

Within lean-startup and product practice the discipline transfers as mechanism with the convenience-trap warning intact: rank the plan's assumptions by consequence × uncertainty × upstream-position, design the cheapest credible falsification of the top entry, run it before sinking downstream work, and treat the awkwardness of a test as irrelevant to its priority. That carries across product categories without translation, with the MVP and presales-pricing instruments as the home-domain ways of running the test.

The notable thing about this entry is that the ranking discipline itself genuinely travels across substrates as mechanism, not analogy — but the lesson it carries belongs to its parent patterns rather than to "RAT" as named, which is the shared-abstract-mechanism case in its strongest form. The same procedure recurs, already named, in field after field: the "killer experiment" / kill-fast pipeline in pharmaceutical R&D (test efficacy in the predictive model before scale-up, to avoid the far larger Phase III failure), red-teaming and pre-mortem rehearsal in military planning (test the assumed adversary response before committing the maneuver), critical-path testing under the technology-readiness-level framework in aerospace, Platt's strong inference in experimental science (run the experiment that maximally discriminates between consequential hypotheses, not the one most likely to "work"), and diligence-budget allocation in venture and credit analysis. These are co-instances of one substrate-independent move — invest the next unit of effort where its result would most change downstream decisions — which the catalog already carries as value_of_information, composed with critical_path (dependency-position), pre_mortem (assumption surfacing), and an explicit assumption inventory. RAT factors cleanly into exactly those parents; its residue is a ranking heuristic codified in lean startup. So when the discipline is wanted in pharma, defense, or research, the honest move is to carry the value-of-information parent (which those fields already instantiate under their own names), not to import "riskiest assumption test" — the cross-field reach is the parent's, and the kill-fast corollary's predictive force is value-of-information's, not a mechanism unique to RAT.

What does not travel, and marks the over-reading boundary, is the named concept itself: outside its lean-startup home "RAT" tends to get borrowed loosely — sometimes flattened to "do early tests," sometimes to "MVP everything" — losing the very ordering discipline (consequence × uncertainty × upstream-position, ranking-over-tooling) that distinguishes it from undifferentiated test-early advice. That loose borrowing is where genuine transfer of the parent mechanism degrades into a slogan. The disciplined statement is therefore that the prioritize-by-value-of-information mechanism recurs everywhere a plan rests on testable assumptions and should be carried as the parent prime, while "riskiest assumption test," with its MVP instruments and startup codification, is the domain-specific name that earns its keep at home and should not be the vehicle for the cross-domain lesson. (See Structural Core vs. Domain Accent.)

Examples

Canonical

The textbook instance is Dropbox in 2008. The plan rested on many assumptions — the sync engine could be built, the pricing would work — but founder Drew Houston judged the riskiest one to be demand: would ordinary people actually want seamless, invisible file synchronization enough to adopt it? Building the full product to find out would have sunk months of engineering into an unverified belief. Instead Houston ran the cheapest credible test of that top-ranked assumption: a roughly three-minute screencast demonstrating the product as if finished, posted to a tech community. The beta waitlist reportedly jumped from about 5,000 to 75,000 overnight — falsifying the "nobody wants this" failure and validating demand before the costly build. The convenience trap would have been to keep coding, the thing the engineers already knew how to do.

Mapped back: Dropbox's coupled build-and-launch plan is the dependency-graphed plan, and buildable / priceable / wanted are the assumption inventory. Demand scored highest on the consequence × uncertainty × upstream ranking — most downstream work rests on it and evidence was weakest. The screencast is the cheapest credible falsification; choosing it over more engineering is the ranking-over-tooling discipline defeating the convenience trap.

Applied / In Practice

AstraZeneca's "5R framework," introduced around 2011 and documented by Cook and colleagues in Nature Reviews Drug Discovery (2014), is a field deployment of the same ordering discipline in pharmaceutical R&D. After a run of costly late-stage failures, the company reframed early development around answering the riskiest assumptions first — right target, right tissue, right patient, right safety, right commercial potential — precisely because a wrong belief about target biology or patient selection, if carried to Phase III, invalidates the largest downstream investment. Rather than advancing compounds on the tests easiest to pass, projects were gated on cheap early evidence for the assumptions whose failure would be most catastrophic downstream. The company subsequently reported a marked improvement in the proportion of projects surviving through the pipeline.

Mapped back: The five R's are the assumption inventory; target and patient biology sit atop the consequence × uncertainty × upstream ranking because they are the most upstream and their failure invalidates the whole Phase III spend — the consequence term and the upstream-position term at their maximum. Early gating on cheap evidence is the cheapest credible falsification, and refusing to advance on merely easy-to-pass tests enacts the ranking-over-tooling discipline against the convenience trap.

Structural Tensions

T1: Auditable ranking versus the noise in its own inputs (the score rests on pre-test guesses). The three-term product — consequence × uncertainty × upstream-position — is what makes RAT auditable, converting a cloud of contingencies into a one-dimensional ordering a team can inspect and defend. But each factor is an estimate made before the testing that would ground it: the team must guess how much downstream work an assumption invalidates and, worse, how weak its own evidence is — a judgment about the very uncertainty the exercise exists to dissolve. The ranking therefore inherits the noise of its inputs; the assumption whose consequence or evidence is hardest to estimate is often the one that most deserves the top slot, yet is exactly where the score is least trustworthy. The crisp product can lend false confidence to an ordering built on soft numbers. Diagnostic: Are the three terms grounded in evidence, or is the ranking's precision masking that its top entry rests on guesses about the very things not yet tested?

T2: Kill-fast falsification versus premature abandonment (the cheapest test can bury a true idea). RAT's action is to run the cheapest credible experiment that could falsify the top assumption and to treat a failure as a signal to pivot before sinking downstream work — the discipline's whole value against sunk cost. But "cheapest" trades fidelity for speed, and a cheap proxy can lie in both directions: a mock-pricing page or a screencast can return a false negative that kills an assumption a truer instrument would have vindicated, or a false positive that validates on a proxy which collapses at scale. The kill-fast virtue and the risk of discarding a real opportunity on a weak signal are the same move seen from two sides — the faster and cheaper the test, the noisier the verdict on which an irreversible pivot may ride. Diagnostic: Is the cheap test a credible proxy for the assumption, or cheap enough that a false negative would kill a viable idea and a false positive would wave through a fatal one?

T3: Ranking-over-tooling versus real capability limits (the awkward top test may be one the team cannot run credibly). The signature discipline is that the ranking, not the team's tooling, governs sequence, so a test's awkwardness or unfamiliarity is declared irrelevant to its priority — this is what defeats the convenience trap. But awkwardness is not always disguised comfort-seeking. A test that sits outside the team's competence may be run badly, and a badly run test of the right assumption can yield a misleading answer worse than none, quietly validating or killing on an artifact of poor execution. The "cheapest credible" clause papers over the conflict: for the genuinely top-ranked assumption the credible test may be neither cheap nor within reach, and insisting on running it anyway can substitute a low-quality answer for an honest "we cannot test this yet." Priority and executability are both real, and they do not always coincide. Diagnostic: Is the top-ranked test being avoided out of mere discomfort, or because the team cannot run it credibly enough to trust the result?

T4: Fixed "riskiest assumption" versus a graph that mutates under testing (the ranking is provisional, not a target). RAT delivers a crisp one-dimensional output — the single assumption atop the ranking — which feels like a fixed target to aim the next experiment at. But the dependency graph it scores is not static: every resolved test collapses one assumption, often spawns new ones, and reweights the upstream-position of the rest, and each pivot the plan makes rewrites the graph wholesale. The "riskiest assumption" is therefore a snapshot that expires the moment it is tested, and the discipline demands continuous re-ranking that its clean single-entry framing can obscure. A team that computes the ranking once and marches down it treats a moving system as a fixed list, and can keep testing yesterday's top assumption while a pivot has already promoted another. Diagnostic: Is this ranking being recomputed as tests resolve and the plan pivots, or is a one-time ordering being followed against a dependency graph that has since changed?

T5: Ranking the inventory versus the assumption no one surfaced (the deadliest belief may not be on the list). Every move RAT makes operates on the assumption inventory — the surfaced set of propositions the plan would fail without. But the ranking can only order what was listed, and the assumption that actually kills a plan is frequently the unarticulated one no one thought to write down: the tacit belief, the unknown unknown, the market or regulatory shift outside the frame. Such an assumption scores nothing because it is absent, so a beautifully ranked, confidently tested inventory can feel like thorough de-risking while the fatal gap sits entirely off the board. RAT's rigor is bounded upstream by the quality of assumption-surfacing (its pre_mortem input), and the more polished the ranking, the easier it is to mistake completeness of the ordering for completeness of the inventory. Diagnostic: Has the inventory been stress-tested for the assumptions no one thought to name, or is the ranking's thoroughness disguising an incomplete list?

T6: Autonomy versus reduction (a lean-startup discipline or the instance of value-of-information). "Riskiest assumption test" is a named, codified discipline with home-domain cargo — the MVP and presales instruments, the convenience-trap warning, the startup framing — and within lean-startup product work it transfers intact as mechanism. But the ordering it enacts, invest the next unit of effort where its result would most change downstream decisions, is not proprietary: it recurs already-named across pharma's kill-fast pipeline, military red-teaming, aerospace critical-path testing, and Platt's strong inference, because all instantiate the parent value_of_information — composed with critical_path (upstream position), pre_mortem (assumption surfacing), and an assumption inventory. RAT factors cleanly into those parents; its residue is a ranking heuristic codified in one field. When the lesson is wanted in pharma or defense, the honest carrier is value-of-information, which those fields already instantiate — not "RAT," whose loose export flattens to "test early" or "MVP everything." Diagnostic: Resolve toward the parent (value-of-information, plus critical_path and pre_mortem) when carrying the lesson to another field; toward the named test when running the ranking-and-experiment discipline inside lean-startup product work in situ.

Structural–Framed Character

The riskiest assumption test sits at the framed-leaning position on the structural–framed spectrum: a codified planning discipline constituted by human decision practice, though short of the framed pole because it embeds a genuinely substrate-neutral decision-theoretic calculus (value-of-information) that runs beneath the startup vocabulary. On evaluative weight it is only mildly framed: RAT is a decision procedure more than a verdict — its output is a ranking and a recommended next experiment, not a conviction — but it does carry a normative edge in "the convenience trap," a named failure mode that indicts running the comfortable test rather than the consequential one, so it prescribes rather than merely describes. On human-practice-bound it is decisively framed: the concept is constituted by the practice of planning under uncertainty and dissolves the instant that practice is removed — its very terms (a team's "comfort zone," a "socially inconvenient" test, "testing for evidence versus testing for comfort," the awkwardness that must be declared irrelevant) presuppose agents making plans, running experiments, and being tempted toward the familiar; strip out the deciding agent and there is no riskiest assumption, only a dependency graph. On institutional origin it is framed: it is a codified discipline of a specific tradition (lean-startup methodology, with its MVP and presales-pricing instruments and its "RAT" branding), an artifact of a management practice rather than a fact of nature. On vocab-travels it is framed at the named level — MVP, pivot, convenience trap, riskiest assumption are pinned to startup practice and flatten to "test early" or "MVP everything" when borrowed loosely — even though the underlying calculus floats free. The one criterion pulling it back toward structure is import-vs-recognize: the entry is emphatic that the ranking discipline genuinely recurs as mechanism, not analogy, already named across pharma's kill-fast pipeline, military red-teaming, aerospace critical-path testing, and Platt's strong inference — but that recognition, it insists, is recognition of the parent, not of RAT.

The portable structural skeleton is value-of-information: invest the next unit of effort where its result would most change downstream decisions, composed with critical_path (upstream position) and pre_mortem (assumption surfacing) over an explicit assumption inventory. That skeleton is precisely what RAT instantiates from those umbrella primes — the entry states it "factors cleanly into exactly those parents," leaving as residue "a ranking heuristic codified in lean startup." The cross-field reach belongs to value-of-information, which pharma, defense, and research already instantiate under their own names; RAT's distinctive cargo — the MVP instruments, the convenience-trap warning, the three-term consequence × uncertainty × upstream ranking as a startup heuristic — stays home. Its character: a practice-constituted, tradition-codified planning discipline, lightly normative through its convenience-trap warning, structural only in the value-of-information calculus it borrows from its umbrella and dresses in lean-startup instruments — framed-leaning, not a free-floating prime.

Structural Core vs. Domain Accent

This section decides why the riskiest assumption test is a domain-specific abstraction and not a prime: a portable value-of-information calculus sits at its core, but the lean-startup codification that makes it the RAT is domain accent that does not lift.

What is skeletal (could lift toward a cross-domain prime). Strip the startup practice and one clean decision rule survives: invest the next unit of effort where its result would most change downstream decisions. An inventory of testable propositions, a scoring by how much each result would reduce expected wasted effort, and a next action aimed at the top. That skeleton is genuinely substrate-portable and is exactly what the catalog carries as value_of_information, composed with critical_path (the upstream-position term), pre_mortem (the assumption-surfacing input), and an explicit assumption inventory. Unusually, the ranking discipline recurs as mechanism, not analogy — already named across pharma's kill-fast pipeline, military red-teaming, aerospace critical-path testing, and Platt's strong inference — which is why RAT "factors cleanly into exactly those parents." But that value-of-information calculus is the core it instantiates, not what makes RAT distinctive; the recurrence is recognition of the parent, not of RAT.

What is domain-bound. Almost all of the concept's distinctive content is lean-startup furniture, and none of it survives extraction as the named thing: the MVP and presales-pricing instruments as the home-domain ways of running the test; the pivot vocabulary; the convenience-trap warning (testing for comfort versus testing for evidence, the engineers who build because building is what they know); and the three-term consequence × uncertainty × upstream ranking codified as a startup heuristic rather than a value-of-information computation over explicit utilities and priors. These are the worked instruments and empirical cases (Dropbox's screencast, AstraZeneca's 5R framework) of lean-startup methodology. The decisive test: strip the codification and what remains is the bare value-of-information move — the residue the entry names is "a ranking heuristic codified in lean startup," with the MVP, the pivot, and the convenience-trap branding all gone. What is left is the parent, not RAT.

Why this does not clear the prime bar. A prime's vocabulary travels and its transfer is recognition of the same mechanism, not analogy. RAT's transfer is bimodal, and this entry is the strongest shared-abstract-mechanism case. Within lean-startup and product practice the discipline travels as full mechanism — the three-term ranking, the cheapest-credible-falsification, the ranking-over-tooling commitment, and the convenience-trap warning mean the same thing across product categories, with the MVP and presales instruments as the home-domain vehicles. Beyond it the named concept degrades: borrowed loosely, "RAT" flattens to "test early" or "MVP everything," losing the very ordering discipline that distinguishes it. What genuinely recurs in pharma, defense, aerospace, and research is the parent, which those fields already instantiate under their own names. And when the substrate-neutral lesson — invest the next effort where its result most changes downstream decisions — is needed cross-field, it is already carried, in more general form, by value_of_information (with critical_path and pre_mortem). The cross-domain reach belongs to those parents; RAT's MVP instruments, convenience-trap warning, and startup codification are the domain accent that stays home in lean-startup product work.

Relationships to Other Abstractions

Local relationship map for Riskiest Assumption TestParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.RiskiestAssumption TestDOMAINPrime abstraction: Assumption — is part ofAssumptionPRIMEPrime abstraction: Prioritization — is part ofPrioritizationPRIMEPrime abstraction: Value of Information — is a decomposition ofValue ofInformationPRIME

Current abstraction Riskiest Assumption Test Domain-specific

Parents (3) — more general patterns this builds on

  • Riskiest Assumption Test is part of Assumption Prime

    The objects ranked and tested are propositions the plan treats as true and on which downstream work depends.

  • Riskiest Assumption Test is part of Prioritization Prime

    RAT ranks competing assumptions by a rule so scarce testing effort goes to the highest-value question first.

  • Riskiest Assumption Test is a decomposition of Value of Information Prime

    Removing lean-startup vocabulary leaves the rule of buying the evidence expected to improve downstream decisions most relative to its cost.

Hierarchy paths (13) — routes to 10 parentless roots

Not to Be Confused With

  • Minimum viable product (MVP). One instrument for running the test in product development — a delivery vehicle that exposes an assumption to the market — not the ordering discipline itself. RAT is about which assumption to test next; the MVP is one way to test it. "MVP everything" loses the consequence × uncertainty × upstream ranking that is RAT's whole content. Tell: is it a vehicle for cheaply exposing an assumption (MVP), or the discipline that ranks which assumption to expose first (RAT)?

  • Value of information (the parent). The formal decision-theoretic calculus — expected reduction in decision loss from an experiment, computed over explicit utilities and priors — that RAT informally instantiates as a ranking heuristic on three estimable terms. RAT is the lean-startup codification; value of information is the substrate-neutral parent that pharma, defense, and research already instantiate under their own names. Tell: is it a computation over explicit utilities and priors (value of information), or the three-term ranking-and-cheapest-test heuristic codified in lean startup (RAT)?

  • Critical-path analysis. The scheduling technique that ranks tasks by time slippage — which delayed task most slips the deadline. RAT ranks by truth slippage — which false belief most invalidates the plan. The upstream-position term borrows critical-path's dependency logic, but the two score different quantities (time vs. consequence-of-being-wrong) even where they coincide. Tell: is the ranking by which task most delays the schedule (critical path), or by which assumption's falsity most wastes downstream work (RAT)?

  • Pre-mortem. The technique of imagining a plan has failed to surface the assumptions that could have killed it. Pre-mortem supplies candidate assumptions — it is RAT's assumption-surfacing input — but it does not rank them or prescribe the next test; the testing-priority commitment is RAT's added move. Tell: is the activity generating the list of what could go wrong (pre-mortem), or ranking that list and picking the cheapest test of the top entry (RAT)?

  • The killer-experiment / strong-inference / red-teaming co-instances (siblings under the parent). The same ordering discipline already named in other fields — pharma's kill-fast "killer experiment," Platt's strong inference in science, military red-teaming, aerospace critical-path testing. These are co-instances of the value-of-information parent, not RAT travelling; each field instantiates the parent under its own name. Tell: outside lean startup, the recurring move belongs to value-of-information (treated in a later section) and to the field's own named version; importing "RAT" there borrows a startup label for a mechanism the field already owns.

Neighborhood in Abstraction Space

Riskiest Assumption Test sits in a sparse region of the domain-specific corpus (70th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Proxy Metrics & Venture Adaptation (13 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-07-12