Arabic dialects present significant challenges to sentiment analysis due to the absence of standardized preprocessing methods, such as stop word removal for Maghrebi Arabic (MA). This paper proposes a novel approach for automatic stop word removal in Maghrebi Arabic (MA) text classification, utilizing two part-of-speech (POS) tagging methods. Using two datasets (AfriSenti and MAC) and two feature extractors (QaRIB and MarBERT), these POS methods were assessed against a predefined stop word list. Furthermore, six neural network architectures were evaluated using 5-fold stratified cross-validation. In the AfriSenti dataset, the combination of automatic stop words removal, MarBERT as a feature extractor, and BiGRU architecture attained the highest accuracy of 75.04% with an F1-score of 62.05%. For the MAC dataset, the best performance was achieved with the same stop words removal method, using Qarib as the feature extractor and BiGRU architecture, resulting in the highest accuracy of 83.31% and an F1-score of 79.28%. This work demonstrates the effectiveness of the POS tagging approach, which emphasizes insightful words and mitigates the limitations of the manual approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Automatic Stop Words Removal in Maghrebi Arabic Dialect Text Classification Using Part of Speech Tagging

  • Yassir Matrane,
  • Faouzia Benabbou,
  • Zineb Ellaky,
  • Chaimae Zaoui

摘要

Arabic dialects present significant challenges to sentiment analysis due to the absence of standardized preprocessing methods, such as stop word removal for Maghrebi Arabic (MA). This paper proposes a novel approach for automatic stop word removal in Maghrebi Arabic (MA) text classification, utilizing two part-of-speech (POS) tagging methods. Using two datasets (AfriSenti and MAC) and two feature extractors (QaRIB and MarBERT), these POS methods were assessed against a predefined stop word list. Furthermore, six neural network architectures were evaluated using 5-fold stratified cross-validation. In the AfriSenti dataset, the combination of automatic stop words removal, MarBERT as a feature extractor, and BiGRU architecture attained the highest accuracy of 75.04% with an F1-score of 62.05%. For the MAC dataset, the best performance was achieved with the same stop words removal method, using Qarib as the feature extractor and BiGRU architecture, resulting in the highest accuracy of 83.31% and an F1-score of 79.28%. This work demonstrates the effectiveness of the POS tagging approach, which emphasizes insightful words and mitigates the limitations of the manual approach.