Deep Learning Approach to Identify and Classify Arabic Verbal Multi-word Expressions
摘要
This paper focuses on the identification and classification of Arabic Verbal Multi-Word Expressions (VMWE), which are combinations of words with a unitary meaning and containing at least one verb. We describe our contribution to annotating the Arabic-annotated-conll17 corpus with VMWEs, following the annotation guide of the PARSEME framework, and using the ChatGPT model to enrich the corpus with sentences that contain more linguistic phenomena. We propose a method for the identification and classification of Arabic VMWEs based on machine learning (Random Forest, Decision trees and Support Vector Machine) and deep learning (Conv1D+BiLSTM) models. The results show that our Random Forest model outperformed others in identification, while the Decision Tree model excelled in classification. Additionally, our deep learning model exhibited significant performance improvement after data enrichment with ChatGPT. These findings underscore the effectiveness of our approach in addressing this complex problem.