A Forced Decoding-based Approach for Enhancing Low-resource ASR
摘要
Multilingual automatic speech recognition represents a crucial research direction in tackling the challenges associated with low-resource scenarios. To effectively incorporate language-specific information in the joint training of a model across multiple languages and leverage linguistic similarities to enhance performance on the target low-resource language, this paper introduces a language similarity evaluation approach based on forced decoding. Specifically, when the target language is specified, the speeches of the source language are decoded into transcription in the target language, and the normalized posterior is utilized as the foundation for evaluating language similarity. Comprehensive experiments and analyses conducted on six low-resource languages reveal that the proposed approach achieves an average word error rate relative reduction of 21.74, 7.68, and 3.45% compared to three widely used benchmark methods, respectively, thereby validating the effectiveness of our approach.