Topic discovery from text data is one of the main subfields of Knowledge Discovery in Databases (KDDs); both aim to extract meaningful patterns from large datasets. KDD typically involves steps such as data cleaning and transformation, while topic discovery (aka topic modeling) methods similarly require text data preprocessing and semantic embedding. Topic modeling also has deep intersections with Natural Language Processing (NLP), and in recent years, methods utilizing pretrained or large language models to discover topics have been proposed. With the rapid development of KDD and NLP, topic modeling has seen significant achievements over the past three decades. Unfortunately, there remains a scarcity of topic models that effectively exploit the temporal aspect, especially the semantic evolution of document streams. Moreover, many contemporary topic models continue to grapple with the issue of noise contamination, particularly in social media data. In this chapter, we will first introduce the mainstream methods for topic modeling, i.e., probabilistic graphical models, nonnegative matrix factorization, and neural networks. Then, we will briefly describe hierarchical and lifelong topic modeling, in addition to the frequently used datasets and evaluation metrics. Finally, we will summarize the challenges in modern topic modeling and topic-level emotion detection from massive short texts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Introduction

  • Yanghui Rao,
  • Qing Li

摘要

Topic discovery from text data is one of the main subfields of Knowledge Discovery in Databases (KDDs); both aim to extract meaningful patterns from large datasets. KDD typically involves steps such as data cleaning and transformation, while topic discovery (aka topic modeling) methods similarly require text data preprocessing and semantic embedding. Topic modeling also has deep intersections with Natural Language Processing (NLP), and in recent years, methods utilizing pretrained or large language models to discover topics have been proposed. With the rapid development of KDD and NLP, topic modeling has seen significant achievements over the past three decades. Unfortunately, there remains a scarcity of topic models that effectively exploit the temporal aspect, especially the semantic evolution of document streams. Moreover, many contemporary topic models continue to grapple with the issue of noise contamination, particularly in social media data. In this chapter, we will first introduce the mainstream methods for topic modeling, i.e., probabilistic graphical models, nonnegative matrix factorization, and neural networks. Then, we will briefly describe hierarchical and lifelong topic modeling, in addition to the frequently used datasets and evaluation metrics. Finally, we will summarize the challenges in modern topic modeling and topic-level emotion detection from massive short texts.