Skip to content

Bitter Lesson

Sutton's historical AI research lesson that general search and learning methods have repeatedly overtaken hand-engineered approaches as they became able to use more computation.

Version
v1 · 2026-10-07 · History
Domain-specific #
13811
Domain group
Applied Sciences & Engineering
Origin domain
Computer Science & Software Engineering
Subdomain
Artificial Intelligence → Computer Science & Software Engineering
Aliases
The Bitter Lesson

Core Idea

The Bitter Lesson is Richard Sutton's historical generalization about AI research: methods built around human understanding of a task can be attractive at the computation available now, while more general search and learning methods have repeatedly become stronger as they found ways to use much more computation. The lesson is “bitter” for researchers who invested in encoding their own domain insight and then watched a less hand-structured route overtake it. Sutton presents this as a recurring history and a guide to research choices, not as a theorem that one method must win on every task or budget.[1]

The comparison needs both sides. A growing compute budget by itself is not the lesson, and a successful learning algorithm by itself does not prove that a particular knowledge-engineered rival plateaued. Sutton points to computer games, speech recognition, and computer vision to argue that research which scales search or learning has repeatedly gained the longer-run advantage.[1]

Structural Signature

Signature: AI task + human-knowledge design + general search or learning design + growing usable computation + observed reversal in a research line → Sutton's Bitter Lesson.[1]

  • Shared AI task and rival designs. The two approaches must address a comparable performance problem. Without that comparison, a story about faster hardware is too broad.[1]
  • Built-in human knowledge. One approach encodes the researcher's understanding as special features, rules, or other task-specific structure. Sutton says this can help in the short run, when available computation is constrained.[1]
  • Search or learning. The other route can spend increased computation searching possibilities or learning behavior or representations. “General” is relative to the hand-built alternative; it does not mean devoid of any human design.[1]
  • Changing usable computation. Falling cost per unit of computation makes a method's capacity to use more of it consequential over research time. More hardware alone cannot rescue an algorithm that cannot put it to productive use.[1]
  • Observed reversal. In Sutton's examples, the search or learning route later outperforms a favored knowledge-heavy direction. This is a historical claim tied to cases, not a universal crossover curve.[1]
  • Research response. The researcher’s effort and attachment to encoded insight explain the name's “bitter” part. They characterize the lesson's interpretation, not the algorithmic performance comparison itself.[1]

What It Is Not

The Bitter Lesson is not a proof that domain knowledge is useless. Sutton allows that building it in can help at the present compute budget; he objects to treating that near-term advantage as the enduring research strategy when search or learning can continue to absorb more computation. Nor is “general” synonymous with an unstructured system: Sutton's computer-vision contrast still names convolution and invariances in modern networks.[1]

It is also not a measured scaling law for an individual model, a prediction of an inevitable date of crossover, or a claim about every field outside AI. The named identity compares research methods and the choices researchers make across an AI history; it is more specific than the broad property of scalability.[1]

Scope of Application

The literal scope is AI research lines in which hand-coded task understanding and compute-using search or learning compete over time. Sutton discusses chess and Go, speech recognition, and computer vision. These cases have different tasks, data, algorithms, and histories. The shared comparison can be used as a research question across them; the essay does not supply one numerical performance curve or universal budget threshold.[1]

Apply the lesson cautiously to a new case: document the two designs, the computation each can productively use, the evaluation conditions, and the actual performance history. If only a small fixed budget is available, or no successful compute-using alternative has been shown, the historical pattern is a hypothesis to test rather than an established result for that case.[1]

Clarity

“More computation” means an approach can turn additional compute into improved search or learning, not merely that the machine is faster. The relevant comparison is between improvement paths under changing resources. A knowledge-rich system and a learned system may both use computation, and the learned one may embody human choices about architecture or objectives. Sutton contrasts where the main task knowledge comes from and how additional compute is converted into performance.[1]

“Bitter” describes the human response to that research reversal. It should not be confused with a negative technical property of search or learning, or read as an instruction to discard current working knowledge regardless of cost.[1]

Manages Complexity

The lesson compresses a recurring method-selection problem into a small set of questions: What task structure did researchers write in? Which search or learning procedure could discover useful structure instead? What additional computation became available, and did the latter procedure use it well enough to change the result? This helps distinguish a short-run engineering win from a research program that improves with resources.[1]

That compression has limits. Sutton's chess and Go histories involve different mixes of search and learning, while his vision account contrasts hand-designed features with deep networks. Naming one lesson does not make these methods or their experimental comparisons identical. The case-specific evidence must still carry the claim.[1]

Abstract Reasoning

When assessing a new AI method, first hold the task and evaluation conditions in view. Compare the human-designed structure and the general search or learning route at a stated resource budget. Ask how each improves if computation grows, and whether the proposed general route can actually exploit that growth. Then seek an observed comparison rather than assuming that falling compute costs force a winner.[1]

If the knowledge-heavy approach remains best under the relevant conditions, the Bitter Lesson does not license rewriting that result. It instead asks whether the apparent lead depends on today's resource limit and whether investing research time in a more scalable route could change the future comparison.[1]

Knowledge Transfer

The comparison transfers within AI from game play to visual perception even though the details do not. In games, Sutton emphasizes search and, for Go, self-play learning; in vision, he contrasts manually designed features and assumptions with convolutional deep learning. The transferable question is where task knowledge is obtained and how additional computation is used, not a single algorithm or proven performance law.[1]

Outside AI, a similar resource-driven reversal might be an analogy. This entry does not treat such analogies as instances of Sutton's named AI research lesson without independent evidence of the rival methods, the resource change, and the outcome.[1]

Examples

Computer chess and Go. Sutton says chess researchers invested in game-specific insight, while the 1997 champion-beating system relied heavily on deep search. He describes Go as a later shift toward effective search and self-play learning. Mapped roles: task → game play; built-in knowledge → researchers' hand-crafted game structure; scalable route → deep search in chess and search with self-play learning in Go; changing computation → the ability to run much more search and training; observed reversal → later scalable approaches displaced favored knowledge-heavy directions in Sutton's account; research response → disappointment at the result. The essay does not establish identical algorithms in both games.[1]

Computer vision. Sutton contrasts earlier edge, generalized-cylinder, and SIFT-style hand-designed approaches with convolutional deep-learning systems. Mapped roles: task → computer vision; built-in knowledge → designed visual features and assumptions; scalable route → learned representations in convolutional networks; changing computation → the essay's broader shift toward methods able to use greater compute; observed reversal → Sutton's field-level account of stronger modern systems; research response → the broader lesson about investment in favored human-designed methods. The essay does not provide a vision-specific measured crossover or imply that convolutional architectures have no human design.[1]

Structural Tensions

Short-run gains versus longer-run capacity to use computation. At today's budget, domain insight can improve a working system. Sutton argues that devoting finite research effort to such encoding can leave less effort for general search or learning methods that may improve more as computation expands. Favoring the immediate method can yield near-term results but risk a plateau; favoring the general route can sacrifice early performance without any guarantee that the hoped-for improvement will occur in this task. The diagnostic is: what evidence shows that each route can use the computation expected over the relevant research horizon? This is Sutton's contingent research-allocation tension, not an optimization law for all algorithms.[1]

Structural–Framed Character

The lesson is framed by an AI research history and Sutton's evaluation of research strategy. Its structure is recognizable across unlike AI tasks: a manually encoded approach, a search or learning competitor, changing usable computation, and a later performance reversal. Its “bitter” interpretation depends on human investment and disappointment. No formal theorem or governing body defines an inevitable outcome; the force of the claim comes from Sutton's reading of repeated cases. The vocabulary of scale travels beyond AI, but the named lesson's evidence and comparison remain in AI method development. Applying it outside AI is an imported analogy unless independent cases establish the same named method comparison and reversal; the present evidence recognizes it literally within AI. Its character: a domain-specific historical rule of thumb whose examples must be tested individually.[1]

Structural Core vs. Domain Accent

The core is a method comparison under changing usable compute: researcher-encoded domain knowledge can win early, while search or learning that productively uses greater computation can later prevail. Remove the competing designs or the resource-responsive reversal and one has a different claim. Chess pieces, Go stones, image features, self-play, SIFT, and a particular network architecture are case accents. Researchers' attachment explains the name but is distinct from the technical performance mechanism.[1]

The broad Prime Scalability names favorable response to added resources; it does not require this AI comparison or Sutton's historical reversal. A model learning curve is a diagnostic trace, not the same research-choice claim. The evidenced cases support a domain-specific identity. Whether the method/resource/reversal skeleton recurs as a portable cross-domain Prime is a future admission question requiring unlike non-AI cases and a strict identity test; no Prime or DAG edge is inferred from this essay alone.[1]

The Bitter Lesson has no broader abstraction in the encyclopedia yet. Scalability is a useful contrast because resource response matters here, but classifying the Bitter Lesson as a kind of Scalability, or treating Scalability as a necessary constituent of it, would take more than topical similarity. Scaling and Scale Dependence is broader still; Machine-Learning Learning Curve describes evaluation of one model across data or optimizer progress; Search Algorithm names one possible method class. None supplies the whole necessary identity of Sutton's historical comparison, so none is counted as a broader abstraction.[1]

Neighborhood in Abstraction Space

Bitter Lesson sits in a sparse region of the domain-specific corpus (93rd percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Cognitive Fixation & Memory Interference (10 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

The essay's observed historical pattern does not promise that search or learning wins every contest. The ability to use more computation is a condition to investigate, not proof of future dominance. Chess search, Go self-play, and visual deep learning are distinct routes. “Bitter” concerns researchers' response, not an objective penalty attached to general methods. A broad statement that systems improve with scale lacks the specific AI research comparison of this entry.[1]

References

[1] Richard S. Sutton, The Bitter Lesson, 13 March 2019. Original author essay, reproduced with permission; especially the opening historical framing, chess and Go paragraphs, computer-vision paragraph, and closing summary. Cited as Sutton's historical argument and examples, not as a controlled proof of universal superiority. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q ↩r ↩s ↩t ↩u ↩v ↩w ↩x ↩y ↩z ↩27 ↩28 ↩29