Building the ArabNER Corpus for Arabic Named Entity Recognition Using ChatGPT and Bard
摘要
The ArabNER corpus marks a significant advancement in the field of Arabic Named Entity Recognition (NER) due to its novel methodology that combines the capabilities of ChatGPT and Bard. In the data generation phase, ChatGPT plays a pivotal role by automatically generating sentences with named entities, ensuring the inclusion of the most recently relevant entities in the corpus. Bard, a powerful language model, is employed for automatic entity annotations. This unique approach greatly reduces the burden of manual linguistic annotation typically required in creating NER corpora for Arabic. Consequently, the ArabNER corpus not only provides a comprehensive dataset for NER research but also sets a precedent for minimizing human intervention in the data annotation process, resulting in a more efficient and up-to-date resource for training and evaluating NER models in Arabic.