Multi-domain Dataset for Moroccan Arabic Dialect Sentiment Analysis in Social Networks
摘要
This chapter is focused on the Sentiment Analysis for Moroccan Arabic dialect text, which is collected from social media platforms. The focus will be to achieve close-to-optimal accuracies for a large dataset of 70,332 entries by using different techniques for feature extraction and machine learning approaches. In doing so, and considering the linguistic complexities in nature and paucity of labeled data for the Arabic dialect context—indeed presenting unique challenges—a host of approaches will be investigated within this chapter to drive sentiment analysis accuracy: TF-IDF, Bag of Words (BOW), Word2Vec, and FastText; also including all Machine learning classification algorithms. One critical part of this study involves the utilization of stemming through Tashaphyne (ArabicLightStemmer), which reduces word variations and improves classification. The findings will contribute to better advancement of sentiment analysis of Moroccan Arabic dialect, offering ample insight useful for social media analytics and linguistic research applications and setting a foundation for future tasks in similar linguistic and computational contexts.