Named Entity Recognition (NER) is a crucial task in Natural Language Processing (NLP), enabling the identification of key entities within texts such as names, locations, and organizations. However, for Arabic, this task is particularly challenging due to its rich morphology, complex syntax, and diverse dialects. This study addresses these challenges by fine-tuning AraBERTv2, a BERT-based model specifically designed for Arabic. We prepare a diverse dataset that covers both Modern Standard Arabic (MSA) and various regional dialects, ensuring broad linguistic coverage. Our results show that the fine-tuned model achieves an F1 Score of 0.859, demonstrating its robustness across different Arabic text forms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing BERT Models for Arabic Named Entity Recognition in a Multi-dialectal Context

  • Bassem Kessentini,
  • Tahar Alimi,
  • Rahma Boujelben,
  • Lamia Hadrich Belguith

摘要

Named Entity Recognition (NER) is a crucial task in Natural Language Processing (NLP), enabling the identification of key entities within texts such as names, locations, and organizations. However, for Arabic, this task is particularly challenging due to its rich morphology, complex syntax, and diverse dialects. This study addresses these challenges by fine-tuning AraBERTv2, a BERT-based model specifically designed for Arabic. We prepare a diverse dataset that covers both Modern Standard Arabic (MSA) and various regional dialects, ensuring broad linguistic coverage. Our results show that the fine-tuned model achieves an F1 Score of 0.859, demonstrating its robustness across different Arabic text forms.