Text classification classifies texts into relevant classes or categories, allowing users to find specific documents on the Internet by searching with the keyword of that targeted class. Both traditional feature-based and transformer-based BERT model for text classifications in the languages have become a very interesting field in 1language processing. Newspapers are the most common source of current events news and information, and online newspaper portals are becoming increasingly popular as people become more Internet dependent. This research aims to classify the Bangla newspaper articles into corresponding classes. Along with text classification, a comparison between the traditional term frequency-inverse document frequency (TF-IDF)-based and bidirectional encoder representations from transformers (BERT)-based model for feature extraction will be brought up in this research. Machine learning algorithms such as decision tree, random forest classifier, support vector machine classifier, and logistic regression model will be implemented for classifying the newspaper articles. A couple of scenarios will be presented and contrasted based on the most accurate prediction, and the maximum accuracy among all instances is 87.96%, which is derived from the logistic regression model and for the transformer-based feature extractor model; hence, the BERT-based model showed superior to the conventional feature extractor methodology.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing TF-IDF and BERT Approach for Bangla Text Classification Using Transformer-Based Embedding for Newspaper Sentiment Classification

  • Samia Binta Hassan,
  • Mohammad Ashiqur Noor,
  • Shamim Forhad,
  • Jarin Tasnim,
  • Abdul Hasib Siddique

摘要

Text classification classifies texts into relevant classes or categories, allowing users to find specific documents on the Internet by searching with the keyword of that targeted class. Both traditional feature-based and transformer-based BERT model for text classifications in the languages have become a very interesting field in 1language processing. Newspapers are the most common source of current events news and information, and online newspaper portals are becoming increasingly popular as people become more Internet dependent. This research aims to classify the Bangla newspaper articles into corresponding classes. Along with text classification, a comparison between the traditional term frequency-inverse document frequency (TF-IDF)-based and bidirectional encoder representations from transformers (BERT)-based model for feature extraction will be brought up in this research. Machine learning algorithms such as decision tree, random forest classifier, support vector machine classifier, and logistic regression model will be implemented for classifying the newspaper articles. A couple of scenarios will be presented and contrasted based on the most accurate prediction, and the maximum accuracy among all instances is 87.96%, which is derived from the logistic regression model and for the transformer-based feature extractor model; hence, the BERT-based model showed superior to the conventional feature extractor methodology.