Twitter-Sentiment Analysis of Moroccan Diabetic: A Comparison Study
摘要
Sentiment analysis of Moroccan diabetic based on social media (ex. Twitter) discussions give coaches necessary knowledge to understand deeply the life style of diabetic people so they can promote adequate recommendations capable of ameliorating the current state of patients. Due to size of data generated in social media, automatization is the only manner to improve the performance of coaches in real time. However, two major challenges persist: dealing with imbalanced datasets and effectively balancing the different sentiments expressed, which can introduce errors into sentiment classification models. Unfortunately, imbalances within tweet corpora and the complexity of feature extraction are at the root of these difficulties, resulting in datasets characterized by excessively high dimensions. Consequently, these factors have a significant impact on the accuracy and efficiency of sentiment analysis performed on these data. This research aims to further the analysis of tweets dealing specifically with diabetes in Morocco, focusing on their content, sentiment, and reach. The designed strategy unfolds in five sequential steps: (a) Data acquisition from Twitter platforms and manual labeling, (b) Utilizing three different representation technique for feature extraction, (c) Utilizing five different variants of oversampling technique, and (d) Using five widely recognized classifiers to classify tweets. The effectiveness of this approach was juxtaposed with that of the classical system, which deals directly with large imbalanced tweet datasets. The comparison was made on the basis of measures such as recall, precision, F1 score, and CPU time. The results show that the proposed system achieves highly accurate sentiment analysis within a reasonable CPU time and outperforms the conventional system in these performance measures.