Sentiment analysis is a crucial area of study in natural language processing that involves the automated interpretation of emotions in textual data. Its effectiveness relies heavily on robust preprocessing techniques that convert raw text into structured inputs suitable for machine learning models. In recent years, ensemble learning methods that combine multiple-based models have been developed to enhance the performance of machine learning models. While many studies have evaluated the effectiveness of different text preprocessing techniques on various machine learning models, the impact of these techniques on ensemble models has not yet been fully explored. Hence, this study performed an experimental analysis of the effectiveness of various preprocessing methods such as lowercasing, stop words removal, punctuation removal, lemmatization, stemming, tokenization and padding sequences on an ensemble model of convolutional neural networks, long short-term memory networks and gated recurrent units for sentiment analysis. The proposed ensemble model based on bagging achieved an accuracy of 0.906 and an F1-score of 0.907 for sentiment analysis on the IMDb dataset using lowercasing, lemmatization, stemming, tokenization and padding sequences. The findings highlight the crucial role of preprocessing in optimizing sentiment analysis models and guiding the selection of preprocessing techniques for sentiment analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Experimental Comparison of the Effect of Text Preprocessing for Sentiment Analysis Using Ensemble Model

  • Hammada Ballali Ebbi,
  • Mohammed Talha Sajidhusein Vasanwala,
  • Seena Joseph

摘要

Sentiment analysis is a crucial area of study in natural language processing that involves the automated interpretation of emotions in textual data. Its effectiveness relies heavily on robust preprocessing techniques that convert raw text into structured inputs suitable for machine learning models. In recent years, ensemble learning methods that combine multiple-based models have been developed to enhance the performance of machine learning models. While many studies have evaluated the effectiveness of different text preprocessing techniques on various machine learning models, the impact of these techniques on ensemble models has not yet been fully explored. Hence, this study performed an experimental analysis of the effectiveness of various preprocessing methods such as lowercasing, stop words removal, punctuation removal, lemmatization, stemming, tokenization and padding sequences on an ensemble model of convolutional neural networks, long short-term memory networks and gated recurrent units for sentiment analysis. The proposed ensemble model based on bagging achieved an accuracy of 0.906 and an F1-score of 0.907 for sentiment analysis on the IMDb dataset using lowercasing, lemmatization, stemming, tokenization and padding sequences. The findings highlight the crucial role of preprocessing in optimizing sentiment analysis models and guiding the selection of preprocessing techniques for sentiment analysis.