错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BERT for Arabic NLP Applications: Pretraining and Finetuning MSA and Arabic Dialects

  • Chaimae Azroumahli,
  • Yacine Elyounoussi,
  • Hassan Badir

摘要

In recent practices, the BERT model that utilizes contextual word Embedding with transfer learning has arisen as a popular state-of-the-art deep learning model. It improved the performance of several Natural Language Processing (NLP) Applications [1]. In this paper, following the effectiveness that these models demonstrated, we use the advantages of training Arabic Transformer-based representational language models to create three Arabic NLP applications for two Arabic varieties; MSA and Arabic Dialects. We build an Arabic representational language model using BERT as the Transformer-based training model [2]. Then we compare the resulting model to the pre-trained multi-lingual models. This step is accomplished by building multiple Arabic NLP applications and then evaluating they are evaluating their performances. Our system had an accuracy of 0.91 on the NER task, 0.89 on the document classification application, and 0.87 on the sentiment Analysis application. This work proved that using a language-specific model outperforms the trained multilingual models on several NLP applications.