Exploring the Efficacy of Text Data Augmentation in Arabic Text Emotion Intensity Classification
摘要
This study examines how data augmentation (DA) enhance the efficacy of machine learning (ML) classifiers in Arabic Emotion Classification (AEC). The classifier's reliability was improved by increasing the dataset using DA strategies to overcome labelled Arabic emotion dataset limitations. In this paper, we applied Random Insertion and Back-translation to augment the dataset. The complete evaluation of classification results is conducted by comparing several classifiers, such as Random Forests (RF), Decision Trees (DT), Support Vector Machines (SVM), Stochastic Gradient Descent (SGD) and logistic regression (LR). The ML models are assessed using the 2018-Ei-OC-Ar-anger Arabic emotion dataset. The proposed method indicates that when datasets are imbalanced, the model performance increase on augmented data. The findings demonstrate that the suggested method enhanced AEC. The classifiers achieved improvements ranging from +89.6 to +168.8% accuracy. The result indicates that text DA has significantly impacted the performance of imbalanced data related to Arabic emotion classification. Hence, the experiment shows that DA enhances the performance of Natural Language Processing (NLP) models.