错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Are Unsupervised Text Classification Techniques Sufficient for Categorizing Short Texts like Product Names?

  • Priya Mishra

摘要

This research paper presents a comparative analysis of unsupervised text classification techniques, specifically Latent Dirichlet Allocation (LDA) and BERTopic. The study employs an empirical approach, combining both quantitative and qualitative analysis, to investigate the effectiveness of these methods in categorizing text without relying on labelled training data, making them suitable for scenarios with limited or unavailable annotated datasets. The evaluation is performed on a dataset, consisting of transaction data from the automobile manufacturing sector. Performance metrics such as the coherence Cv, coherence CUMass, MoverScore, and perplexity are assessed to compare LDA and BERTopic. Additionally, the study analyzes the quality and interpretability of discovered topics, providing valuable insights into the classification outcomes. Significantly, the innovation of this study lies in the application of these topic modeling techniques to extremely short documents, effectively addressing the challenge of categorizing such concise texts. These findings contribute to advancing text analysis techniques and hold practical implications for applications with scarce labelled training data. The innovation offers a practical solution for industries dealing with short and unlabeled texts, including e-commerce, customer reviews and social media analysis. Researchers exploring the unsupervised text classification methods will also find value in the insights provided into the performance and interpretability of LDA and BERTopic.