In the last decade, the number of texts and complex documents has increased exponentially, requiring a deeper knowledge of machine learning techniques for carefully classifying texts. The most fundamental task in natural language processing is Text classification. In natural language processing, many machine-learning algorithms have achieved excellence. The success of these approaches depends on their ability to understand nonlinear relationships and complicated models in data. In recent years, research in this area has seen an upsurge because of the unprecedented success rate of Deep Learning. Various datasets, evaluation metrics, and methods are suggested in the literature, so an up-to-date and comprehensive review is needed. In text classification tasks, Deep Learning-based models have outperformed traditional machine learning-based approaches in several areas (question answering, sentiment analysis, topic labeling, natural language inference, news classification, and named entity recognition). This paper discusses these tasks in detail, covering benchmark datasets and technical developments. This study eliminates the gap by providing an overview of the state of the art from 2014 to 2022, and the target is all tasks and models, from traditional to Deep Learning. Text classification tasks are thoroughly analyzed from over 70 articles, discussing technical contributions, strengths, and commonalities. This study compares various approaches, listing evaluation criteria along with their pros and cons. However, due to recent developments, emerging trends, and the rapid evolution of models, this study is influenced by specific datasets and metrics, leading to particular results. Future work should include more recent research, broader datasets, further comparisons, and the exploration of new technologies in text classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text Classification: A Comprehensive Survey from Traditional Approaches to Deep Learning Methods

  • Ihab L. Hussein Alsammak,
  • Wasan H. Itwee,
  • Moamin A. Mahmoud,
  • Noor Islam Jasim

摘要

In the last decade, the number of texts and complex documents has increased exponentially, requiring a deeper knowledge of machine learning techniques for carefully classifying texts. The most fundamental task in natural language processing is Text classification. In natural language processing, many machine-learning algorithms have achieved excellence. The success of these approaches depends on their ability to understand nonlinear relationships and complicated models in data. In recent years, research in this area has seen an upsurge because of the unprecedented success rate of Deep Learning. Various datasets, evaluation metrics, and methods are suggested in the literature, so an up-to-date and comprehensive review is needed. In text classification tasks, Deep Learning-based models have outperformed traditional machine learning-based approaches in several areas (question answering, sentiment analysis, topic labeling, natural language inference, news classification, and named entity recognition). This paper discusses these tasks in detail, covering benchmark datasets and technical developments. This study eliminates the gap by providing an overview of the state of the art from 2014 to 2022, and the target is all tasks and models, from traditional to Deep Learning. Text classification tasks are thoroughly analyzed from over 70 articles, discussing technical contributions, strengths, and commonalities. This study compares various approaches, listing evaluation criteria along with their pros and cons. However, due to recent developments, emerging trends, and the rapid evolution of models, this study is influenced by specific datasets and metrics, leading to particular results. Future work should include more recent research, broader datasets, further comparisons, and the exploration of new technologies in text classification.