Enhancing Named Entity Recognition in Modern Standard Arabic via Fine-Grained Part-of-Speech Tags
摘要
Arabic, a morphologically rich language, poses unique challenges for named entity recognition (NER) due to its lack of capitalization, complex word forms, and significant ambiguity. This study investigates the impact of incorporating fine-grained POS tags in ANER, demonstrating an F1 score improvement from 74.2% to 87.5% across four datasets with increasing annotation complexity. These findings highlight the importance of linguistic context in improving ANER and suggest broader applications in machine translation, sentiment analysis, and other NLP tasks.