The quality of word representation is essential for achieving high performance in various Natural Language Processing (NLP) tasks. This study investigates the impact of pre-trained word embeddings on sentiment classification of Arabic text using deep learning (DL) models. The study compares the performance of two pre-trained word embedding models namely, AraVec and FastText, generated at both tweet and word levels. Additionally, we propose a hybrid word embeddings approach combining both embedding methods. Sentiment classification was performed using multiple DL models, including Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), Long Short-Term Memory networks (LSTM), Bidirectional LSTM (Bi-LSTM), and Gated Recurrent Units (GRU) for comparative study. The experimental results revealed that hybrid embeddings outperformed individual embedding methods across all classification models. Notably, the Bi-LSTM model with hybrid embedding at the word-level achieved the highest accuracy of 94.52% and an F1-score of 93.21%, demonstrating the effectiveness of combining AraVec and FastText word embeddings. AraVec consistently outperformed FastText due to its training on a massive corpus of Arabic tweets, making it more suitable for the tweet dataset. The findings highlight the importance of selecting appropriate embeddings based on the model architecture and the nature of the task, providing valuable insights for future research in Arabic sentiment analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Arabic Sentiment Analysis Leveraging Hybrid Word Embeddings with Deep Learning Techniques

  • Abdulrahman Alharbi,
  • Nabin Sharma,
  • Farookh Hussain

摘要

The quality of word representation is essential for achieving high performance in various Natural Language Processing (NLP) tasks. This study investigates the impact of pre-trained word embeddings on sentiment classification of Arabic text using deep learning (DL) models. The study compares the performance of two pre-trained word embedding models namely, AraVec and FastText, generated at both tweet and word levels. Additionally, we propose a hybrid word embeddings approach combining both embedding methods. Sentiment classification was performed using multiple DL models, including Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), Long Short-Term Memory networks (LSTM), Bidirectional LSTM (Bi-LSTM), and Gated Recurrent Units (GRU) for comparative study. The experimental results revealed that hybrid embeddings outperformed individual embedding methods across all classification models. Notably, the Bi-LSTM model with hybrid embedding at the word-level achieved the highest accuracy of 94.52% and an F1-score of 93.21%, demonstrating the effectiveness of combining AraVec and FastText word embeddings. AraVec consistently outperformed FastText due to its training on a massive corpus of Arabic tweets, making it more suitable for the tweet dataset. The findings highlight the importance of selecting appropriate embeddings based on the model architecture and the nature of the task, providing valuable insights for future research in Arabic sentiment analysis.