PashtoEmo: Enhancing Text-Based Emotion Analysis in the Pashto Language Through Dataset Creation
摘要
This paper presents the comprehensive PashtoEmo dataset for emotion analysis in the Pashto language. PashtoEmo contains 8016 text instances from \(\mathbb {X}\) social media platform covering several topics, including politics, women’s rights, social justice, culture, sports, and education. The texts are annotated with six emotional categories and an additional ‘Other’ class. We tested PashtoEmo for emotion analysis with several machine learning, deep learning, and transformer models. Even the best model, XLM-RoBERTa-large, achieved only a 76.85% F1 score and a 78.11% accuracy. This indicates that the data set must be extended with adequate resources in Pashto language text-based applications.