Automatic Accent Identification Using Less Data: a Shift from Global to Segmental Accent
摘要
Accentedness is a prominent feature of foreign language learning. While humans have a remarkable capacity to adapt their perception to accents, they remain a hard challenge to the robustness of automatic speech recognition (ASR). In particular, the necessity to use large non-native annotated datasets for model training remains a crucial issue. This paper investigates the possibility of reducing the data need of these systems by using targeted training datasets, focusing on the most challenging rather than on all non-native phonemes. Specifically, the study examines whether training data for ASR, and accent identification systems in particular, could focus not on global but on segmental accent. Segmental accent refers to the deviations in pronouncing specific phonemes, while global accent captures the extent to which a non-native speaker is perceived to differ from a native one. An accent identification problem was formulated, where models were trained on two types of data: full words (i.e., global accent) versus isolated difficult vowels (i.e., segmental accent), both uttered by natives and non-natives. Two novel highly controlled and professionally annotated datasets were used for that purpose. Throughout experiments, a transfer learning approach using pretrained deep residual neural networks was applied, with subsequent comparison to a baseline support vector machine. Results showed that although word-based classification yielded better accuracy, the dataset consisting of isolated vowels could account for much of the accent, when used with both methods (up to 80%). Applications of this approach and the possibility of using smaller, but more representative datasets are discussed.