Wav2vec-AD: Acoustic Unit Discovery Module-Integrated, Self-Supervised Contrastive Pre-training Approach for Speech Recognition
摘要
An effective speech recognition model necessitates an ample supply of labeled data for supervised training. However, this proposition poses a monumental challenge for low-resource languages in terms of constructing a speech recognition system with high precision. In this paper, we propose a novel pre-training strategy for contrastive learning by fusing the acoustic unit discovery module with Wav2vec 2.0, herein referred to as Wav2vec-AD. This strategy, for the first time in speech contrastive learning, enables controlled negative sample selection via the acoustic unit discovery module, thereby augmenting the model’s representational learning capability. Furthermore, we conduct a thorough analysis regarding the selection of negative samples in different situations to enhance the speech representation learned by the model, optimizing its efficacy in downstream tasks. In the low-resource case, compared to the baseline Wav2vec 2.0, Wav2vec-AD achieves absolute word error rate (WER) improvements of 1.55% and 1.46% respectively on the development-clean and test-clean subsets of LibriSpeech. Moreover, absolute WER improvements of 0.63% and 4.21% were realized in Arabic and Turkish language datasets, respectively.