Automatic summarization¶
Automatic production of a shorter text retaining salient information from source material.
Core Idea¶
Automatic summarization is the computational production of a shorter representation that preserves information judged important for a task, audience, or query. The input may be one document, many documents, an image collection, audio, video, or mixed data; the output may be text, key phrases, representative images, key frames, or selected segments. A summarizer must therefore define both a selection objective—importance, coverage, novelty, relevance, chronology, or user need—and a compression constraint. Shortening alone is insufficient: deleting arbitrary content produces a smaller object but not a defensible summary.
Extractive systems choose units already present in the source, such as sentences, phrases, frames, or shots. They trade expressive compression for traceability and lower fabrication risk, but can yield redundancy or incoherent transitions. Abstractive systems generate a new representation by paraphrasing, combining, and reorganizing source content. They can state a compact synthesis that no source unit expresses alone, but must manage factual consistency, attribution, and unsupported inference. Hybrid and human-aided workflows use machines to propose salient material and people to revise scope or accuracy. Generic summarization aims at broad coverage; query-focused summarization conditions importance on a specified information need.
Evaluation is relational: a summary is assessed against its source, task, length budget, and users. Fluency does not prove faithfulness, and lexical overlap can undervalue a correct abstraction while rewarding copied but unimportant text. Multi-document inputs add contradiction, temporal update, source diversity, and redundancy problems. Video synopsis that synthesizes new composite frames is likewise different from selecting key shots. The abstraction is the constrained computational transformation from a larger information object to a smaller task-representative one, with explicit losses and preservation criteria.
How would you explain it like I'm…
The Short-Version Maker
Computer-Made Summaries
Task-Guided Automatic Summaries
Structural Signature¶
Sig role-phrases:
- the source information object — one or more documents, images, recordings, videos, or mixed-media collections
- the audience-or-query context — the user need that determines what counts as important
- the preservation objective — coverage, relevance, novelty, chronology, representativeness, or another declared content priority
- the compression budget — an explicit constraint on length, duration, frames, or output size
- the selection-or-generation mechanism — extraction of existing units, abstraction into new expressions, or a hybrid workflow
- the redundancy-and-coherence control — organization that avoids repeated content and preserves intelligible relations among retained points
- the faithfulness constraint — support of output claims by the source with correct attribution and contradiction handling
- the shorter representation — text, key phrases, frames, segments, or composite output satisfying the task
- the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone
What It Is Not¶
- Not arbitrary shortening. A defensible summary preserves information judged important under a task, audience, query, and length constraint.
- Not necessarily text-to-text. Inputs and outputs can include documents, audio, video, images, key frames, phrases, or mixed media.
- Not one extractive method. Extraction selects source units, whereas abstraction generates a new synthesis with different traceability and fabrication risks.
- Not proven accurate by fluency. A polished output can misstate, invent, or misattribute information from the source.
- Not evaluated by lexical overlap alone. Copying can preserve unimportant wording, while a faithful abstraction can use different language.
- Not context-free compression. Generic and query-focused summaries preserve different things, and multi-source tasks must handle contradiction, redundancy, chronology, and provenance.
- Not lossless representation. The method should make its preservation objective and accepted omissions explicit rather than pretending the smaller object contains everything.
Scope of Application¶
Automatic summarization applies when a computational system must produce a shorter information object under explicit audience, task, query, and length constraints while preserving selected source content.
- Single-document text. Extractive, abstractive, or hybrid systems reduce articles, reports, transcripts, and records for defined users.
- Multi-document synthesis. Redundant and conflicting sources require provenance, temporal ordering, and explicit handling of disagreement.
- Query-focused output. Selection is conditioned on a user's information need rather than general salience alone.
- Conversation and meeting summaries. Decisions, participants, uncertainty, and action items must remain attributable to the source interaction.
- Audio, image, and video. Temporal or visual segments can be selected or described under modality-specific fidelity criteria.
- Evaluation and operations. Coverage, relevance, redundancy, coherence, factual consistency, citation, currency, compression, and task utility must be measured.
- Applicability boundary. Arbitrary shortening, keyword deletion, and fluent source-ungrounded generation are not valid summarization; high-stakes uses require source access, auditability, and proportionate human review.
Clarity¶
Automatic summarization makes shortening answerable to an information objective and a compression constraint. It distinguishes extractive selection from abstractive generation, single-document from multi-document synthesis, and generic coverage from query- or audience-focused relevance. The term prevents a shorter output from counting as a summary merely because content was deleted. The sharper evaluation question is which source-supported information must survive, what redundancy and ordering constraints apply, and how faithfulness, coverage, coherence, and length are traded under the intended use.
Manages Complexity¶
Automatic summarization turns a large source collection into an optimization over coverage, relevance, novelty, redundancy, coherence, faithfulness, and length. The analyst specifies the audience or query, information units, compression budget, and error costs rather than judging ‘shorter’ as a single property. Extractive and abstractive branches trade traceability against expressive compression; single- and multi-document settings add different redundancy and contradiction problems. This representation permits evaluation by what important content survives, what unsupported content appears, and how usable the ordered result is, without re-reading the entire source for every system comparison.
Abstract Reasoning¶
Selection move. From source units and an importance or relevance objective, infer which content deserves limited summary capacity. Compression move. Combine or rephrase units only when their source support and distinctions survive; otherwise prefer traceable extraction. Coverage move. Compare the output against required topics and redundancy to infer what information was lost or overrepresented. Faithfulness move. Trace every asserted fact to the input and reject fluent additions unsupported there. Boundary move. A shorter output is not evidence of successful summarization, and evaluation must match the intended audience, query, medium, and acceptable tradeoff among length, coherence, and detail.
Knowledge Transfer¶
Within the home domain. Automatic summarization transfers across news, scientific literature, meetings, legal documents, dialogue, and multimedia when a system selects or generates a shorter representation preserving task-relevant content. Source grounding, compression, salience, redundancy, coherence, and evaluation retain operational roles. Beyond the home domain (C — computational instrument). It applies literally wherever an input representation and summary objective are defined. Its boundary is over-reading: brevity does not guarantee factuality, coverage, neutrality, or suitability for a user; reference metrics do not exhaust quality. Human memory, abstraction, and institutional reporting may summarize, but are not automatic summarization unless an algorithm performs the transformation.
Examples¶
Canonical¶
Given a news article, an extractive summarizer may score sentences for relevance to a query, penalize redundancy, and select a short ordered subset. If the article explains a policy decision, the summary should preserve the decision, responsible actor, effective date, and central reason while omitting repeated background. Selecting only the first sentence may meet a length budget but fail the preservation objective; selecting several near-duplicate sentences wastes the budget and harms coherence. An abstractive system may rephrase instead, but then every generated claim must remain grounded in the source. A good summary is relational: its adequacy depends on the intended audience and task, not on shortness alone.
Mapped back: The article is the source information object, the reader's need the audience-or-query context, and required facts the preservation objective. Sentence scoring is the selection-or-generation mechanism, length is the compression budget, diversity is the redundancy-and-coherence control, and grounding is the faithfulness constraint.
Applied / In Practice¶
A meeting assistant converts a transcript into decisions, unresolved questions, and assigned actions. Speaker turns and discussion are the source, while the audience needs operational follow-up rather than a miniature transcript. The system groups repeated discussion, preserves who owns each action and any deadline explicitly stated, and links summary items back to transcript spans. Participants review the result before distribution, correcting uncertain names and rejecting inferred commitments that were never agreed. Evaluation therefore checks factual support, coverage of consequential items, and usefulness for the next meeting, not only overlap with one human reference summary.
Mapped back: Transcript and agenda supply the source information object and audience-or-query context. Decisions and actions are the preservation objective within the compression budget; grouping supplies the redundancy-and-coherence control, transcript links enforce the faithfulness constraint, and participant review performs the relational evaluation of the shorter representation.
Structural Tensions¶
T1 — Identity versus admissible variation. Automatic summarization must remain recognizable across legitimate variants. Admissible variation is bounded by this condition: Extractive, abstractive, or hybrid systems reduce articles, reports, transcripts, and records for defined users. The stable element is expressed by this invariant: Automatic production of a shorter text retaining salient information from source material. Treating every surface change as a new abstraction fragments the identity, while allowing a change to the constitutive relation produces a false positive.
Diagnostic: After the proposed variation, can an analyst still establish this invariant: Automatic production of a shorter text retaining salient information from source material?
T2 — Recognition versus proxy. The domain needs observable or inferential evidence for Automatic summarization, but the evidence is not automatically the identity. The working recognition rule is: the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone. A familiar indicator can occur without the defining relation, and the relation can persist when a customary detector is unavailable.
Diagnostic: Does the evidence establish the defining claim—Automatic production of a shorter text retaining salient information from source material—or only a correlated sign?
T3 — Definition versus operational judgment. A compact definition aids reuse, whereas actual classification in natural-language processing can require expert decisions about boundary conditions, measurements, conventions, or exceptions. Extractive systems choose units already present in the source, such as sentences, phrases, frames, or shots. The definition must constrain those judgments without pretending that every admissible case can be recognized from a label alone.
Diagnostic: Which observation would make a competent practitioner reject the classification under the stated definition?
T4 — Scope versus overextension. Automatic summarization has a genuine habitat in which extractive, abstractive, or hybrid systems reduce articles, reports, transcripts, and records for defined users. Yet Arbitrary shortening, keyword deletion, and fluent source-ungrounded generation are not valid summarization; high-stakes uses require source access, auditability, and proportionate human review. A useful application map therefore has to be broad enough to cover recurring practice and narrow enough to exclude merely topical or metaphorical occurrences.
Diagnostic: Can the claimed application fill the same carrier and relation roles, or has only the name traveled?
T5 — Transfer versus domain accent. Knowledge about Automatic summarization can travel within its home domain, and some structural lessons may travel farther. Automatic summarization transfers across news, scientific literature, meetings, legal documents, dialogue, and multimedia when a system selects or generates a shorter representation preserving task-relevant content. What transfers must be separated from the specialist vocabulary, warrant, and closure conditions that remain anchored in natural-language processing.
Diagnostic: Is the receiving case a literal instance of Automatic summarization, a co-instance of Representation, or only an analogy?
T6 — Autonomy versus reduction. Automatic summarization is a strict specialization of Representation, but the edge does not erase the domain differentia. The broader node supplies only the necessary structural relation; natural-language processing supplies the carrier, warrant, boundary, and exception conditions expressed by this identity: Automatic production of a shorter text retaining salient information from source material. The entry is over-split if those conditions add no discriminating work and under-specified if the parent alone is used for cases that require them.
Diagnostic: Can a domain expert use the added conditions to distinguish Automatic summarization from another case that equally instantiates Representation?
Structural–Framed Character¶
Automatic summarization is mixed: structurally specifiable but materially dependent on its disciplinary frame. Its structural side consists of the carrier the source information object — one or more documents, images, recordings, videos, or mixed-media collections and the constitutive relation Automatic production of a shorter text retaining salient information from source material. Its framed side comes from natural-language processing, which fixes what the terms denote, what counts as evidence, and when a qualification or exception defeats the classification.
Across the principal tests, the entry is not merely a free-floating pattern. Evaluative weight: the identity can be stated descriptively even when its use has practical or normative consequences. Practice dependence: the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone. Institutional stabilization: disciplinary conventions may stabilize the name and test without necessarily creating every underlying event or relation. Vocabulary portability: the invariant is Automatic production of a shorter text retaining salient information from source material. Import versus recognition: an outside case qualifies literally only if the same typed roles and collapse condition are available; otherwise the comparison is analogical.
The reusable remainder is Representation under a reviewed subsumption relation. That node preserves the necessary cross-domain organization after the natural-language processing-specific carrier, evidence, and exceptions are removed. Automatic summarization remains autonomous because its recognition and collapse conditions distinguish cases that the parent alone leaves together.
Structural Core vs. Domain Accent¶
What is skeletal. The portable skeleton is a typed carrier organized by a constitutive relation, an invariant, a recognition test, and a collapse condition. Here the carrier is the source information object — one or more documents, images, recordings, videos, or mixed-media collections. The decisive relation is Automatic production of a shorter text retaining salient information from source material, which also states the controlling invariant at this level. Stripped of specialist nouns, this organization is represented by Representation.
What is domain-bound. natural-language processing supplies the actual objects or agents, admissible transformations, units or conventions, standards of warrant, and named exceptions. In this case, recognition requires evidence for the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone. Admissible variation is bounded by the condition that extractive, abstractive, or hybrid systems reduce articles, reports, transcripts, and records for defined users, and the classification collapses when a defensible summary preserves information judged important under a task, audience, query, and length constraint. These are constitutive differentia, not illustrative decoration.
Why it remains a domain-specific node. The reviewed DAG relation is subsumption to Representation. Outside natural-language processing, the parent captures only the reusable structural remainder. The specialist name remains literal only where the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone can be established under the domain's standards of warrant.
Instantiates / Related Primes¶
This entry is a kind of Representation.
- Immediate parent — Representation (subsumption). Automatic summarization is a domain-specific kind of Representation: Automatic production of a shorter text retaining salient information from source material. The parent supplies the necessary broader identity—Model complex ideas.—while the candidate adds the source-domain carrier, recognition rule, and failure conditions. The defining source account begins: Automatic summarization is the computational production of a shorter representation that preserves information judged important for a task, audience, or query.
- Nearest catalog surface declined — Automatic item generation. Its rematch score was 0.200104. Retrieval proximity did not establish synonymy or parentage; the carrier, invariant, and collapse condition remain different.
- Related reasoning operations. Evidence, comparison, boundary testing, and representation can support a case without becoming additional DAG parents.
Relationships to Other Abstractions¶
Current abstraction Automatic summarization Domain-specific
Parents (1) — more general patterns this builds on
-
Automatic summarization is a kind of Representation Prime
Automatic summarization is a domain-specific kind of Representation: Automatic production of a shorter text retaining salient information from source material.The parent supplies the necessary broader identity—Model complex ideas.—while the candidate adds the source-domain carrier, recognition rule, and failure conditions. The defining source account begins: Automatic summarization is the computational production of a shorter representation that preserves information judged important for a task, audience, or query.
Hierarchy path (1) — routes to 1 parentless root
- Automatic summarization → Representation → Abstraction
Neighborhood in Abstraction Space¶
Automatic summarization sits in a moderately populated region (57th percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Unclustered & Miscellaneous (2551 abstractions)
Nearest neighbors
- Memory (Rhetorical Canon) — 0.87
- Temporal Distinctiveness — 0.85
- Mimesis — 0.85
- Format Relation — 0.85
- Mandela Effect — 0.85
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Representation. This is the reviewed immediate parent or structural prerequisite, not a synonym. Tell: retain Automatic summarization only when the domain-specific relation
Automatic production of a shorter text retaining salient information from source material.and its source-domain warrant are established; otherwise route the case to Representation. -
Latent Semantic Analysis. This is the closest catalog retrieval surface, not an accepted synonym or parent. Tell: Ask which entry's carrier, invariant, and collapse test the case actually satisfies; shared vocabulary or a score of 0.703626 is insufficient.
-
Not arbitrary shortening. A defensible summary preserves information judged important under a task, audience, query, and length constraint. Tell: Require the positive recognition condition that the relational evaluation — assessment against source, task, users, and budget rather than fluency or lexical overlap alone.
-
Not necessarily text-to-text. Inputs and outputs can include documents, audio, video, images, key frames, phrases, or mixed media. Tell: Replace the familiar surface feature and test whether automatic production of a shorter text retaining salient information from source material.
-
A detector, representation, or consequence. A method may reveal Automatic summarization, a notation may describe it, and an outcome may follow from it without any of those being identical to the abstraction. Tell: Would the defining relation remain if the present detector, notation, or downstream effect changed?
-
A metaphorical transfer. A case outside the home domain may resemble the structure while lacking its native role types and standards of warrant. Tell: If only the general organization survives, route the comparison to Representation rather than treating it as another Automatic summarization instance.
References¶
- Frozen Wikipedia revision: https://en.wikipedia.org/wiki/Automatic_summarization (revision 1347086022).
- DOI: https://doi.org/10.1109/tvcg.2019.2948611
- DOI: https://doi.org/10.1109/mcg.2011.89
- DOI: https://doi.org/10.1109/CVPR.2012.6247852
- DOI: https://doi.org/10.1109/TIP.2016.2615289
- DOI: https://doi.org/10.1016/j.ins.2017.12.020
- DOI: https://doi.org/10.1007/978-3-319-66939-7_19
- DOI: https://doi.org/10.1023/A:1009976227802
- DOI: https://doi.org/10.3103/S0005105510030027
- Supporting reference preserved in the packet: https://www.wiley.com/en-gb/Automatic+Text+Summarization-p-9781848216686
- Supporting reference preserved in the packet: https://www.proquest.com/docview/1986931333
- Supporting reference preserved in the packet: https://books.google.com/books?id=O0fNBQAAQBAJ&q=video+surveillance+summarization&pg=PA81
- Supporting reference preserved in the packet: https://research-information.bris.ac.uk/files/111433536/Ioannis_Pitas_Multimodal_Stereoscopic_Movie_Summarization_Conforming_to_Narrative_Characteristics.pdf
- Supporting reference preserved in the packet: https://www.sciencedirect.com/science/article/abs/pii/S0020025517311398
- Supporting reference preserved in the packet: http://ai.googleblog.com/2022/03/auto-generated-summaries-in-google-docs.html
- Supporting reference preserved in the packet: https://www.dummies.com/education/language-arts/speed-reading/how-to-skim-text/
- Supporting reference preserved in the packet: http://acl.ldc.upenn.edu/acl2004/emnlp/pdf/Mihalcea.pdf
The frozen Wikipedia revision is discovery provenance. The cited source set was reviewed for identity, formal or operational relation, and scope. The encyclopedia's structural synthesis is bounded to those claims; URL transport failure alone was not treated as substantive contradiction.