Being humans makes us have the ability to analyze and understand the meanings of the sentences clearly and detect where these sentences come from or what the sentiment polarity of these sentences is. So, as humans, we can figure out the spoken dialects, the sentiment polarity or any other classification tasks but it has to be said that this ability cannot be obtained by machines easily. The present article is aimed at exploring how machines can understand dialects and sentiment polarity in Arabic language using the deep learning model AraBERT. It is a pretrained model which is used for Arabic language. The identification tasks are not just binary identification tasks but we have used more than two labels in the process of detecting. We have evaluated the work of AraBERT model with Arabic dialects using three different corpora and proved that we can get a better accuracy of 90% for binary classification and a better accuracy of 87% for multiclass classification. The most important part of this research is that we also used a dataset from Arabic online newspapers written in standard form of Arabic and could achieve of 95% accuracy for multiclass classification using AraBERT while we could achieve of 91% accuracy using normal machine learning models for the same dataset. We are making the used corpora and the source code freely available for research purposes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating AraBERT Model in Detecting Dialects and Sentiment Polarity in Arabic Language

  • Nabeel Abdulrazaq Yaseen,
  • Javad Vahidi,
  • Sukaina Abdulhussein Abdullah,
  • Hani Akram Mahfoud

摘要

Being humans makes us have the ability to analyze and understand the meanings of the sentences clearly and detect where these sentences come from or what the sentiment polarity of these sentences is. So, as humans, we can figure out the spoken dialects, the sentiment polarity or any other classification tasks but it has to be said that this ability cannot be obtained by machines easily. The present article is aimed at exploring how machines can understand dialects and sentiment polarity in Arabic language using the deep learning model AraBERT. It is a pretrained model which is used for Arabic language. The identification tasks are not just binary identification tasks but we have used more than two labels in the process of detecting. We have evaluated the work of AraBERT model with Arabic dialects using three different corpora and proved that we can get a better accuracy of 90% for binary classification and a better accuracy of 87% for multiclass classification. The most important part of this research is that we also used a dataset from Arabic online newspapers written in standard form of Arabic and could achieve of 95% accuracy for multiclass classification using AraBERT while we could achieve of 91% accuracy using normal machine learning models for the same dataset. We are making the used corpora and the source code freely available for research purposes.