<p>Despite being spoken by more than 400 million people worldwide, Arabic language datasets remain considerably scarce compared to other well-resourced languages. The high cost of manual annotation is frequently cited as the primary obstacle. While Semi-Supervised Self-Learning (SSSL) techniques offer popular alternatives for pseudo-labeling unlabeled data, they are computationally expensive and provide limited explainability. Lexicon-based approaches are more time-efficient but suffer from accuracy limitations due to their dependence on lexicon quality. This paper introduces a novel approach that leverages Explainable AI (XAI) techniques to automatically generate high-quality sentiment lexicons for Arabic text pseudo-labeling. The work makes three key contributions: (A) We propose an enhanced algorithm that extends existing XAI-based lexicon generation methods by incorporating computed parameters specifically designed for Arabic text characteristics and introducing novel lexicon balancing techniques. (B) We conduct comprehensive evaluation on an Arabic sentiment analysis task using three different classifier architectures: traditional machine learning, deep learning, and transformer-based models. (C) We demonstrate competitive performance (F1 = 0.82) compared to the SSSL ensemble approach (F1 = 0.88) while offering significantly improved computational efficiency, explicit interpretability, and better cross-domain generalization. The proposed algorithm provides a favorable trade-off between performance, efficiency, and transparency, making it particularly suitable for applications requiring explainable AI or resource-constrained environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An auto-generated XAI-based lexicon approach for pseudo-labeling datasets in sentiment analysis tasks

  • Ahmed El-Sayed,
  • Amr Magdy,
  • Moustafa Abd El Haliem,
  • Ziad Mohamed,
  • Zeyad Ahmed,
  • Antoine Abdelmalak,
  • Omar Khairat,
  • Mohamad Mamdouh,
  • Shaimaa Lazem

摘要

Despite being spoken by more than 400 million people worldwide, Arabic language datasets remain considerably scarce compared to other well-resourced languages. The high cost of manual annotation is frequently cited as the primary obstacle. While Semi-Supervised Self-Learning (SSSL) techniques offer popular alternatives for pseudo-labeling unlabeled data, they are computationally expensive and provide limited explainability. Lexicon-based approaches are more time-efficient but suffer from accuracy limitations due to their dependence on lexicon quality. This paper introduces a novel approach that leverages Explainable AI (XAI) techniques to automatically generate high-quality sentiment lexicons for Arabic text pseudo-labeling. The work makes three key contributions: (A) We propose an enhanced algorithm that extends existing XAI-based lexicon generation methods by incorporating computed parameters specifically designed for Arabic text characteristics and introducing novel lexicon balancing techniques. (B) We conduct comprehensive evaluation on an Arabic sentiment analysis task using three different classifier architectures: traditional machine learning, deep learning, and transformer-based models. (C) We demonstrate competitive performance (F1 = 0.82) compared to the SSSL ensemble approach (F1 = 0.88) while offering significantly improved computational efficiency, explicit interpretability, and better cross-domain generalization. The proposed algorithm provides a favorable trade-off between performance, efficiency, and transparency, making it particularly suitable for applications requiring explainable AI or resource-constrained environments.