Improving mispronunciation detection and diagnosis for non- native learners of the Arabic language
摘要
Mispronunciation detection and diagnosis (MDD) is a core component of computer-assisted pronunciation training (CAPT), which aims to provide opportunities for second language learners (L2) to learn and practice their speaking skills. Arabic is one of the most widespread languages in the world, with more than 422 million speakers. It is the language of the Holy Quran, which increases the importance of learning Arabic. Most existing Arabic MDD systems focus on learning-based techniques rather than state-of-the-art deep-learning methods. Most existing Arabic MDD systems are primarily relying on traditional learning techniques. However, integrating Transformer-based algorithms into the system is crucial, as it can significantly enhance the accuracy, efficiency, and overall performance of the Arabic MDD systems. This paper introduces an Arabic MDD system using transformer-based techniques for non-native learners of spoken Arabic language to enhance their learning of Arabic and help non-native speakers practice their pronunciation skills. The study focuses on detecting mispronunciation phonemes and analyzing these pronunciation errors, by identifying the type of each error, whether it was insertion, deletion, or substitution error. To train the MDD system, we constructed a speech dataset for native and non-native Arabic speakers. The performance evaluation was obtained based on the phoneme recognition and MDD performance, and we achieved a phoneme error rate of 3.1%, 80.8%, and 91.3% for diagnosis accuracy and detection accuracy, respectively. Additionally, we conducted a human perceptual test to assess human proficiency in detecting pronunciation errors and compare their evaluations with automatic verification. The automatic verification surpassed human verification, achieving a detection accuracy of 97.0% and a diagnosis accuracy of 80.4%.