Social networks have become a major source for the expression and dissemination of information and news in various languages. Users can express their opinions and sentiments towards different topics through written comments in dialectal languages. The objective of this work is to build a multi-class sentiment analysis model for Tunisian dialect comments scraped from social media websites, based on a lexical ontology and deep learning models. First, we created the TDCOR_Train, a multi-domain Tunisian Arabic dialect dataset. Subsequently, we developed the TDSO lexical ontology, which will be projected to automatically annotate each comment as positive, negative, or neutral. Further-more, we trained four deep learning models, namely Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Bi-LSTM, and RCNN (RNN + CNN) on the annotated dataset. Both RCNN and Bi-LSTM produced the best results, achieving accuracy rates of 81.27 and 82.01. To assess the performance of our sentiment analysis method, we fine-tuned the bert-base-arabertv02 transfer transformer which returned an accuracy rate of 81.17. Moreover, all models exceeded the accuracy rate of 85.00 when evaluated on two test corpora, the first is TSAC_New, and the second is TSAC_New_TDSO, which is TSAC_New dataset auto-annotated by the TDSO lexical ontology.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Opinion Analysis Based on a Sentiment Lexical Ontology and Deep Learning Models: Tunisian Dialect Case

  • Tahar Alimi,
  • Rahma Boujelben,
  • Lamia Hadrich Belguith

摘要

Social networks have become a major source for the expression and dissemination of information and news in various languages. Users can express their opinions and sentiments towards different topics through written comments in dialectal languages. The objective of this work is to build a multi-class sentiment analysis model for Tunisian dialect comments scraped from social media websites, based on a lexical ontology and deep learning models. First, we created the TDCOR_Train, a multi-domain Tunisian Arabic dialect dataset. Subsequently, we developed the TDSO lexical ontology, which will be projected to automatically annotate each comment as positive, negative, or neutral. Further-more, we trained four deep learning models, namely Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Bi-LSTM, and RCNN (RNN + CNN) on the annotated dataset. Both RCNN and Bi-LSTM produced the best results, achieving accuracy rates of 81.27 and 82.01. To assess the performance of our sentiment analysis method, we fine-tuned the bert-base-arabertv02 transfer transformer which returned an accuracy rate of 81.17. Moreover, all models exceeded the accuracy rate of 85.00 when evaluated on two test corpora, the first is TSAC_New, and the second is TSAC_New_TDSO, which is TSAC_New dataset auto-annotated by the TDSO lexical ontology.