Skip to content

Text mining

Text mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text.

Version
v1 · 2026-09-28 · History
Domain-specific #
12506
Domain group
Interdisciplinary & Synthetic
Origin domain
Data Science & Analytics
Subdomains
Text Analytics, Natural Language Processing → Data Science & Analytics

Core Idea

Text mining is treated here as the recurring text analytics identity summarized by this source-grounded definition: Text mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text. Text mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text. It involves "the discovery by computer of new, previously unknown information, by automatically extracting information from different written resources." Written resources may include websites, books, emails, reviews, and articles.

Scope of Application

  • Scientific literature mining and academic applications. The Text Analysis Portal for Research (TAPoR), currently housed at the University of Alberta, is a scholarly project to catalogue text analysis applications and create a gateway for researchers new to.

  • Text analytics. The latter term is now used more frequently in business settings while "text mining" is used in some of the earliest application areas, dating to the 1980s, notably life-sciences research and.

  • Applications. In business, applications are used to support competitive intelligence and automated ad placement, among numerous other activities.

  • Security applications. Many text mining software packages are marketed for security applications, especially monitoring and analysis of online plain text sources such as Internet news, blogs, etc. for national security purposes.

  • Situation in the United Kingdom. However, owing to the restriction of the Information Society Directive (2001), the UK exception only allows content mining for non-commercial purposes.

Clarity

A clear use of Text mining names the carrier, the operative relation, and the conditions under which the source treats the identity as present. The minimal definition is Text mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text.

Manages Complexity

Text mining compresses multiple text analytics details into a stable diagnostic relation. The source shows both the central mechanism—now, through use of a semantic web, text mining can find content based on meaning and context (rather than just by a specific word).—and the practical consequence—text mining is being used by large media companies, such as the Tribune Company, to clarify information and to provide readers with greater.

Abstract Reasoning

  1. Type the carrier. Identify the text analytics entities to which the claim applies.
  2. State the relation. Use the source-grounded identity: Text mining, text data mining (TDM) or text analytics is the process of deriving high-quality information from text.
  3. Check operation and conditions. Text mining methods and software is also being researched and developed by major firms, including IBM and Microsoft, to further automate the mining and analysis processes, and by different firms working in the area of search and indexing in general as a way.

Knowledge Transfer

Within the home domain. Knowledge about Text mining transfers literally when a new case preserves the same carrier type, relation, and recognition test. The Text Analysis Portal for Research (TAPoR), currently housed at the University of Alberta, is a scholarly project to catalogue text analysis applications and create a gateway for researchers new to the practice. The latter term is now used more frequently in business settings while "text mining" is used in some of the earliest application areas, dating to the 1980s, notably life-sciences research and government intelligence. Beyond the home domain.

Neighborhood in Abstraction Space

Text mining sits in a sparse region of the domain-specific corpus (64th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Unclustered & Miscellaneous (2551 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08