<p>End-to-end speech recognition has benefited from large-scale labeled corpora to achieve great success, and limited labeled data hardly meets performance requirements in specific scenarios. Self-supervised learning provides a compelling solution for addressing this issue. However, existing self-supervised methods cannot balance context-semantic relationships and differentiated information between speech features, which are crucial for ASR. To address this issue, we propose a novel Dual-Stream Self-Supervised Learning Network (DSSLNet) to combine the complementary advantages of both parties. Concretely, the dual-stream structure consists of a reconstruction prediction module and a contrastive prediction module in parallel, where reconstruction prediction is jointly trained with contrastive prediction and as an auxiliary task of the latter. Furthermore, a novel GRU feature fusion module is also designed for fusing speech representations, which adaptively explores the latent structure of speech through a parameter learning strategy. Our DSSLNet is first pre-trained on Multi-Audio (600h) and Librispeech (960h), then fine-tuned on Aishell-1, HKUST and subsets of Librispeech for the ASR task. Experiment results show that our DSSLNet achieves state-of-the-art compared to other advanced works while achieving comparable accuracy in limited labeled data scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSSLNet: dual-stream self-supervised learning network for end-to-end speech recognition

  • Boyang Lyu,
  • Chunxiao Fan,
  • Yue Ming,
  • Nannan Hu

摘要

End-to-end speech recognition has benefited from large-scale labeled corpora to achieve great success, and limited labeled data hardly meets performance requirements in specific scenarios. Self-supervised learning provides a compelling solution for addressing this issue. However, existing self-supervised methods cannot balance context-semantic relationships and differentiated information between speech features, which are crucial for ASR. To address this issue, we propose a novel Dual-Stream Self-Supervised Learning Network (DSSLNet) to combine the complementary advantages of both parties. Concretely, the dual-stream structure consists of a reconstruction prediction module and a contrastive prediction module in parallel, where reconstruction prediction is jointly trained with contrastive prediction and as an auxiliary task of the latter. Furthermore, a novel GRU feature fusion module is also designed for fusing speech representations, which adaptively explores the latent structure of speech through a parameter learning strategy. Our DSSLNet is first pre-trained on Multi-Audio (600h) and Librispeech (960h), then fine-tuned on Aishell-1, HKUST and subsets of Librispeech for the ASR task. Experiment results show that our DSSLNet achieves state-of-the-art compared to other advanced works while achieving comparable accuracy in limited labeled data scenarios.