Named Entity Recognition (NER) plays a vital role in extracting structured information from unstructured text, especially on social media platforms such as Twitter. Deep neural networks are effective for a variety of natural language processing tasks; nevertheless, they often require large sets of annotated data sets to outperform simpler models. This data may not be sufficiently diverse or available, and collecting and annotating them can be a time-consuming and costly process. The aim of this paper is to study the impact of data augmentation on the Named Entity Recognition (NER) task on a text issue from social media, using pre-trained models in Arabic. Two scenarios were set up to evaluate the models, starting with a scenario without data augmentation, in which the MARBERT v2 model surpassed the other models with an F1 score of 67.4%. Then, the second scenario with data augmentation, in which the Arabert v0.2 twitter model obtained the best F1 score with 72.16%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Arabic NER in Social Media: Performance Analysis of Pre-trained Models and Data Augmentation

  • Brahim Ait Benali,
  • Soukaina Mihi,
  • Toufik Datsi,
  • Azeddine Elmajidi

摘要

Named Entity Recognition (NER) plays a vital role in extracting structured information from unstructured text, especially on social media platforms such as Twitter. Deep neural networks are effective for a variety of natural language processing tasks; nevertheless, they often require large sets of annotated data sets to outperform simpler models. This data may not be sufficiently diverse or available, and collecting and annotating them can be a time-consuming and costly process. The aim of this paper is to study the impact of data augmentation on the Named Entity Recognition (NER) task on a text issue from social media, using pre-trained models in Arabic. Two scenarios were set up to evaluate the models, starting with a scenario without data augmentation, in which the MARBERT v2 model surpassed the other models with an F1 score of 67.4%. Then, the second scenario with data augmentation, in which the Arabert v0.2 twitter model obtained the best F1 score with 72.16%.