Assessing BERT Models for Arabic Named Entity Recognition in a Multi-dialectal Context
摘要
Named Entity Recognition (NER) is a crucial task in Natural Language Processing (NLP), enabling the identification of key entities within texts such as names, locations, and organizations. However, for Arabic, this task is particularly challenging due to its rich morphology, complex syntax, and diverse dialects. This study addresses these challenges by fine-tuning AraBERTv2, a BERT-based model specifically designed for Arabic. We prepare a diverse dataset that covers both Modern Standard Arabic (MSA) and various regional dialects, ensuring broad linguistic coverage. Our results show that the fine-tuned model achieves an F1 Score of 0.859, demonstrating its robustness across different Arabic text forms.