<p>This paper introduces AlgVec, a suite of word embedding models trained specifically for the Algerian dialect, a linguistically rich but under-resourced variety of Arabic. The embeddings are derived from a large corpus of user-generated content on the social media platform X (formerly Twitter), covering both Arabic script and Arabizi. The Algerian dialect presents unique challenges for Natural Language Processing (NLP), including informal grammatical structures, non-standardized spelling, and frequent code-switching. To address these issues, we compile and preprocess a corpus of more than 32 million tokens and train multiple embedding models using Word2Vec (Skip-gram and Continuous Bag of Words (CBOW)) as well as FastText architectures. We evaluate AlgVec through intrinsic tasks such as word similarity, nearest-neighbor retrieval, and a linguistically grounded DiaLex-style benchmark adapted for Arabic dialects, as well as through a downstream sentiment analysis task. Furthermore, we explore a combined embedding approach that integrates AlgVec with the transformer-based MARBERT model for sentiment analysis, using Support Vector Machine (SVM), Convolutional Neural Network (CNN), and Long Short-Term Memory (LSTM) classifiers. Our experiments show that AlgVec and the combined model outperform widely used Modern Standard Arabic (MSA) embeddings on dialect-specific benchmarks. The complete set of embeddings is publicly released to support future research in Arabic dialect NLP.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AlgVec: A word embedding model for the algerian dialect in arabic and arabizi

  • Lamia Ouchene,
  • Sadik Bessou

摘要

This paper introduces AlgVec, a suite of word embedding models trained specifically for the Algerian dialect, a linguistically rich but under-resourced variety of Arabic. The embeddings are derived from a large corpus of user-generated content on the social media platform X (formerly Twitter), covering both Arabic script and Arabizi. The Algerian dialect presents unique challenges for Natural Language Processing (NLP), including informal grammatical structures, non-standardized spelling, and frequent code-switching. To address these issues, we compile and preprocess a corpus of more than 32 million tokens and train multiple embedding models using Word2Vec (Skip-gram and Continuous Bag of Words (CBOW)) as well as FastText architectures. We evaluate AlgVec through intrinsic tasks such as word similarity, nearest-neighbor retrieval, and a linguistically grounded DiaLex-style benchmark adapted for Arabic dialects, as well as through a downstream sentiment analysis task. Furthermore, we explore a combined embedding approach that integrates AlgVec with the transformer-based MARBERT model for sentiment analysis, using Support Vector Machine (SVM), Convolutional Neural Network (CNN), and Long Short-Term Memory (LSTM) classifiers. Our experiments show that AlgVec and the combined model outperform widely used Modern Standard Arabic (MSA) embeddings on dialect-specific benchmarks. The complete set of embeddings is publicly released to support future research in Arabic dialect NLP.