错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating K-Means Clustering and Levenshtein Distance and K-Nearest Neighbor Algorithms for Enhanced Arabic Sentiment Analysis

  • Ghaleb Al-Gaphari,
  • Salah AL-Hagree,
  • Hamzah A. Alsayadi

摘要

Arabic sentiment analysis (ASA) is a challenging field due to the complexity of the Arabic language. Although there have been some studies on ASA, the number of such studies is relatively limited compared to those conducted on English or other Latin languages, indicating a research gap. In this paper, we propose a new approach to ASA based on the comments dataset of users of mobile applications available on the Google play store. The proposed approach combines the K-Nearest Neighbor (K-NN) and K-Means Clustering (K-MC) algorithms with the Levenshtein distance (LD) algorithm for data preprocessing and feature extraction. A number of experiments were conducted to evaluate the performance of these algorithms in the context of ASA using mobile application reviews. The results of this study reveal that by integrating the K-NN, K-MC, and LD algorithms, we achieved superior performance compared to both the standalone K-NN algorithm and the combination of K-NN with LD. The integrated approach yielded impressive outcomes, with an accuracy of 84.12%, recall of 68.08%, precision of 85.33%, and F-score of 75.74%. Furthermore, it led to enhancements of 1.01% in accuracy, 1.78% in recall, 0.23% in precision, and 1.21% in F-score. The proposed study contributes to the field of ASA by proposing a novel approach that improves the performance of sentiment analysis on Arabic scripts in the context of mobile application reviews.