错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Impact of Preprocessing Techniques Towards Word Embedding

  • Mustazzihim Suhaidi,
  • Rabiah Abdul Kadir,
  • Sabrina Tiun

摘要

In this study, we analyze the performance of various pre-processing methods and classification algorithms on health-related tweet data on Twitter. The data set consists of a number of different pre-processing methods, such as Z-score Scaling, Min-max Scaling, Decimal Scaling, Log Transformation, Percentage Scaling, and Log2 Scaling, as well as two main classification algorithms: Naive Bayes and Logistic Regression. The results of the analysis show that the pre-processing method has a significant effect on the performance of the classification algorithm. Z-score Scaling emerges as a stable option and provides good accuracy for both algorithms. However, Min-max Scaling is more suitable for Logistic Regression than Naive Bayes. In addition, Logistic Regression tends to provide higher accuracy in some pre-processing methods. We also suggest further exploration to understand how this pre-processing method might apply to different types of data and its impact on other classification algorithms. In addition, the selection of alternative models, hyperparameter optimization, and data enrichment are areas that can be improved to obtain better classification results. This study underscores the importance of careful pre-processing and selection of appropriate pre-processing methods in applying classification algorithms to text data. The results and recommendations in this study can be a guide for researchers and practitioners in making better decisions in the classification analysis of Twitter data about health.