错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sentiment Analysis of Libyan Middle Region Using Machine Learning with TF-IDF and N-grams

  • Abdullah Habberrih,
  • Mustafa Ali Abuzaraida

摘要

Arabic dialects are commonly used on social media platforms by Arabic speakers to express their opinions and connect with each other. However, due to the lack of standardized rules or grammars, analyzing Arabic dialects with NLP tools can be more challenging than standard Arabic. Moreover, the poems domain within Arabic dialects is considered to be more challenging than other domains due to structural differences between poems and regular expressions. This study investigates the use of TF-IDF and N-grams with Lemmatization techniques to develop machine learning classifiers for sentiment analysis in the domain of poems within the Libyan dialect, specifically the Libyan Middle Region dialect. Three experiments were conducted using ML classifiers, namely, SVM, NB, and LR. The first experiment explored classifiers’ performance using .TF-IDF with Unigrams, whereas the second experiment investigated the use of TF-IDF with Trigrams, and the third experiment examined the impact of combining Unigrams and Trigrams on the classifiers’ performance. The experimental results indicate that utilizing Unigrams with TF-IDF can enhance classifier performance, whereas Trigrams with TF-IDF can have a negative effect on the classifiers’ performance. Notably, SVM achieved the highest accuracy of 69.04% in the first experiment, while LR achieved the highest accuracy of 58.60% in the second experiment, and SVM achieved an accuracy of 68.92% in the third experiment. Furthermore, in terms of precision, LR in the first experiment outperformed the other classifiers across all experiments with 70.49%.