Explainable AI Driven Deep Learning for Accurate L1 Identification from L2 Speech
摘要
The automatic identification of speakers’ native language (L1) through the analysis of non-native language (L2) speech has garnered considerable attention in the fields of speech recognition and language assessment. In this study, an extensive investigation is presented, focusing on feature engineering mechanisms within deep learning architectures for L1 identification from L2 speech. The efficacy of convolutional neural networks (CNNs), residual neural networks (ResNets) and vision transformer (ViT) models, is explored in distinguishing L2 speech originating from different L1 backgrounds. To enhance interpretability, Explainable AI (XAI) techniques, specifically Shapley Additive Explanations (SHAP), are employed, offering insights into the learned features and decision-making processes. Experimental results highlight the effectiveness of the examined architectures and the value of XAI in facilitating transparent analysis. This research contributes to the advancement of knowledge in the areas of language identification and deep learning, while also paving the way for further investigations in feature engineering and the integration of XAI techniques. The NISP dataset comprises English speeches and their native language from 345 native speakers of Hindi, Tamil, Telugu, Malayalam, and Kannada. The training and testing accuracy pairs of the CNN, ResNet18, and ViT models are (88.3%, 98.8%, and 99.8%) and (89.6%, 96.8%, and 97.87%) respectively.