Lottery ticket hypothesis¶
In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
Core Idea¶
Lottery ticket hypothesis is treated here as the recurring machine-learning pruning identity summarized by this source-grounded definition: In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. The original statement of the hypothesis is: A randomly-initialized, dense neural network contains a subnetwork that is initialized such that—when trained in isolation—it can match the test accuracy of the original network after training for at most the same number of iterations. Magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative).
Reset surviving weights to initialization: use m \odot \theta_0 as the starting point. Train the masked network f!\left(x ; m \odot \theta_0\right) with the same optimizer, schedule, and data. Declare "winning ticket" if it achieves comparable test accuracy in \leq j iterations.
For Lottery ticket hypothesis, the abstraction is narrower than the article's general subject matter: a positive case must preserve In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. Retaining only the name, a familiar example, or a downstream effect is insufficient. The specialist roles and tests remain anchored in machine-learning pruning, which is why this identity is domain-specific rather than prime.
Structural Signature¶
Sig role-phrases:
- Defining carrier — In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
- Constitutive relation — Magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative).
- Operating condition — However, after training, these lottery tickets can be discovered by the pruning algorithm.
- Recognition evidence — Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning.
- Admissible variation — A similar result has been proven for the special case of convolutional neural networks.
- Characteristic consequence — Reset surviving weights to initialization: use m \odot \theta_0 as the starting point.
- Failure boundary — Train the masked network f!\left(x ; m \odot \theta_0\right) with the same optimizer, schedule, and data.
What It Is Not¶
- Not the whole field of machine-learning pruning. The node requires the specific identity stated by In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
- Not an over-broad reading. Such networks are a priori difficult to find, since before training, one does not know which of the exponentially many subnetwork would be "lottery tickets".
- Not an over-broad reading. However, after training, these lottery tickets can be discovered by the pruning algorithm.
- Not an over-broad reading. It was found that if instead of m \odot \theta_0 , they re-sampled a different random initialization \theta_0' , and used m \odot \theta_0' instead, the trained f!\left(x ; m \odot \theta_0'\right) would perform much worse.
- Not automatically Lottery mathematics. Retrieval proximity does not establish equivalence; the two identities must be compared by carrier, operation, and failure boundary.
Scope of Application¶
Lottery ticket hypothesis applies literally inside machine-learning pruning wherever the source-defined carrier and relation can be established. Its documented habitats include:
- Documented setting. It was found that if instead of m \odot \theta_0 , they re-sampled a different random initialization \theta_0' , and used m \odot \theta_0' instead, the trained f!\left(x ; m \odot \theta_0'\right) would perform much worse.
- Subsequent work. Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning.
- Subsequent work. A similar result has been proven for the special case of convolutional neural networks.
- Documented setting. In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
- Documented setting. Magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative).
- Documented setting. Reset surviving weights to initialization: use m \odot \theta_0 as the starting point.
Outside machine-learning pruning, the name should be retained only when these same operational conditions survive; otherwise the comparison belongs to the broader parent Theory or should be marked as analogy.
Clarity¶
A clear use of Lottery ticket hypothesis names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. The strongest recognition evidence in the frozen account is: Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning. A report should distinguish that evidence from a proxy, consequence, or common implementation. It should also state the qualification Such networks are a priori difficult to find, since before training, one does not know which of the exponentially many subnetwork would be "lottery tickets". so that a reader can reproduce the classification rather than infer it from topical resemblance.
Manages Complexity¶
Lottery ticket hypothesis compresses multiple machine-learning pruning details into a stable diagnostic relation. The source shows both the central mechanism—magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative).—and the practical consequence—reset surviving weights to initialization: use m \odot \theta_0 as the starting point. This compression makes cases comparable while leaving parameters, conventions, exceptions, and evidential quality explicit. It is lossy by design: local history and implementation details may be omitted only when they do not alter the defining relation.
Abstract Reasoning¶
- Type the carrier. Identify the machine-learning pruning entities to which the claim applies.
- State the relation. Use the source-grounded identity: In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network.
- Check operation and conditions. However, after training, these lottery tickets can be discovered by the pruning algorithm.
- Demand recognition evidence. Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning.
- Test variation. Change an implementation or setting while preserving a similar result has been proven for the special case of convolutional neural networks.
- Run the collapse test. Remove the defining operation; if the label still seems equally apt, only a topic or correlate was retained.
- Reduce cautiously. When the specialist conditions cannot be carried, route the residual comparison to Theory.
Knowledge Transfer¶
Within the home domain. Knowledge about Lottery ticket hypothesis transfers literally when a new case preserves the same carrier type, relation, and recognition test. It was found that if instead of m \odot \theta_0 , they re-sampled a different random initialization \theta_0' , and used m \odot \theta_0' instead, the trained f!\left(x ; m \odot \theta_0'\right) would perform much worse. Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning.
Beyond the home domain. No canonical parent is asserted for Lottery ticket hypothesis. An outside case receives the specialist name only when the same typed roles and rejection conditions can be filled literally; otherwise the comparison remains an analogy pending later graph densification.
Examples¶
Canonical¶
A similar result has been proven for the special case of convolutional neural networks. This case is canonical because it supplies a concrete carrier and lets the defining relation be checked rather than merely named.
Mapped back: carrier → the entities in the documented case; operation → In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network; recognition evidence → Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning
Applied / In Practice¶
Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning. The applied case shows how the identity is used under a second setting or qualification while keeping the same operative relation.
Mapped back: changed setting → Subsequent work; invariant → In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network; boundary → the case exits the class when such networks are a priori difficult to find, since before training, one does not know which of the exponentially many subnetwork would be "lottery tickets"
Structural Tensions¶
T1 — Stable identity versus admissible variation. Such networks are a priori difficult to find, since before training, one does not know which of the exponentially many subnetwork would be "lottery tickets". The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Which changes preserve the defining relation, and which replace it?
T2 — Recognition versus proxy. However, after training, these lottery tickets can be discovered by the pruning algorithm. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Does the cited evidence establish the identity or only a correlated sign?
T3 — Definition versus implementation. It was found that if instead of m \odot \theta_0 , they re-sampled a different random initialization \theta_0' , and used m \odot \theta_0' instead, the trained f!\left(x ; m \odot \theta_0'\right) would perform much worse. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Is the observed implementation constitutive, optional, or merely common?
T4 — Scope versus overextension. Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Can every claimed application fill the same typed roles without metaphor?
T5 — Transfer versus domain accent. In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: Does the receiving case instantiate Lottery ticket hypothesis literally, co-instantiate Theory, or only resemble it?
T6 — Autonomy versus reduction. Magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative). The tension matters because emphasizing only one side either dissolves the identity or overstates what the evidence and domain conventions warrant.
Diagnostic: What does Lottery ticket hypothesis distinguish that the broader parent Theory leaves together?
Structural–Framed Character¶
Lottery ticket hypothesis is mixed or framed-leaning. Its structural side is the repeatable organization summarized by In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. Its framed side is the machine-learning pruning vocabulary that fixes the carrier, evidence, exceptions, and admissible transformations.
Evaluative weight: the identity can be stated descriptively even when applications carry practical stakes. Human-practice dependence: the source-grounded carrier determines whether the relation exists independently or is constituted by a practice. Institutional origin: disciplinary conventions stabilize the name and test. Vocabulary portability: However, after training, these lottery tickets can be discovered by the pruning algorithm. Import versus recognition: literal transfer requires the same mechanism; shape alone is analogy.
Its portable skeleton is Theory. Its character: a recurring specialist identity whose thin organization can be abstracted, while its operational meaning remains domain-bound.
Structural Core vs. Domain Accent¶
What is skeletal. In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. The stable skeleton is the typed relation expressed in that definition and the entry's recognition and collapse tests. The source identifies these operative conditions: In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. Magnitude pruning: build a bitmask m by setting to 0 the lowest-magnitude weights in each layer (one-shot), or repeat train → prune across rounds to reach higher sparsity (iterative). It further constrains recognition and variation through: However, after training, these lottery tickets can be discovered by the pruning algorithm. Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning.
What is domain-bound. machine-learning pruning supplies the operative entities, technical vocabulary, warrants, and exceptions that make Lottery ticket hypothesis literal. Its documented scope includes the condition that It was found that if instead of m \odot \theta0 , they re-sampled a different random initialization \theta0' , and used m \odot \theta0' instead, the trained f!\left(x ; m \odot \theta0'\right) would perform much worse. Another bounded application condition is that Malach et al. proved a stronger version of the hypothesis, namely that a sufficiently overparameterized untuned network will typically contain a subnetwork that is already an approximation to the given goal, even before tuning. These are not decorative examples; they determine which carrier and evidence can fill the abstraction's roles.
Why no parent is asserted. Removing those specialist details does not currently yield one live catalog node that is a necessary genus for every instance. The entry is therefore approved as unparented rather than attached by topical resemblance. Its collapse evidence remains specific—A similar result has been proven for the special case of convolutional neural networks.—and future graph densification may discover a defensible relation only if it preserves that boundary.
Instantiates / Related Primes¶
This entry is a kind of Scientific Hypothesis.
- Approved unparented node. No current live node supplies a defensible necessary genus or structural prerequisite for Lottery ticket hypothesis. The reviewed identity is: In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network. The accelerated suggestion was declined because topical or lexical similarity does not establish hierarchy; the node is admitted without a parent pending later graph densification.
- Related reasoning operations. Evidence, representation, comparison, classification, transformation, or evaluation may participate in particular cases, but participation does not make any one of them a necessary parent of every instance.
Relationships to Other Abstractions¶
Current abstraction Lottery ticket hypothesis Domain-specific
Parents (1) — more general patterns this builds on
-
Lottery ticket hypothesis is a kind of Scientific Hypothesis Domain-specific
It is a machine-learning hypothesis with experimental consequences.It is a machine-learning hypothesis with experimental consequences.
Hierarchy path (1) — routes to 1 parentless root
- Lottery ticket hypothesis → Scientific Hypothesis → Falsifiability
Neighborhood in Abstraction Space¶
Lottery ticket hypothesis sits in a sparse region of the domain-specific corpus (91st percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.
Family — Combinatorial Optimization & Discrete Structures (31 abstractions)
Nearest neighbors
- Lottery mathematics — 0.81
- Gambling and information theory — 0.80
- Lottery paradox — 0.79
- Random compact set — 0.79
- Large width limits of neural networks — 0.79
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Theory. The parent omits the specialist differentia. Tell: Can the case establish In machine learning, the lottery ticket hypothesis is that artificial neural networks with random weights can contain a subnetwork which (entirely by chance) can be tuned to a similar performance as tuning the whole network?
- Lottery mathematics. The combinatorial and probabilistic analysis of lottery drawings, prize tiers, expected returns and apparent coincidences under declared game rules. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- Lottery paradox. The inconsistency between accepting each highly probable claim that an individual lottery ticket will lose and accepting the certain claim that some ticket will win. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- Randomness Test. Challenge a sequence against a specified stochastic null using a pattern-sensitive statistic and calibrated rejection rule, while treating a pass only as failure to detect the tested departures. Tell: Which entry's carrier, operation, and failure condition are satisfied?
- A measurement, proxy, or consequence. Those may provide evidence without being the identity. Tell: Would Lottery ticket hypothesis remain present if the detector or downstream effect changed?
- A metaphorical analogue. A similar shape outside machine-learning pruning lacks the specialist mechanism. Tell: Do the native roles transfer literally, or only the parent Theory?
References¶
- Frozen Wikipedia discovery revision: https://en.wikipedia.org/wiki/Lottery_ticket_hypothesis (revision 1348911377).
- Preserved source candidate: https://hal.science/hal-03548226
The frozen Wikipedia revision is discovery provenance. The retained source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; a thin authority surface is recorded as a nonblocking source-strengthening repair rather than concealed.