<p>BSRNN, a band-split RNN model designed for tasks such as speech enhancement and source separation, has demonstrated outstanding performance in recent research. Meanwhile, contrastive learning, as a self-supervised learning framework, has proven effective in guiding networks to learn speech representations with broad potential across domains. This paper explores integrating contrastive learning techniques into BSRNN to optimize its speech enhancement capabilities. Specifically, we design various contrastive losses that complement the supervised loss to jointly train the model. To improve the shallow acoustic representations generated by BSRNN’s band-split module, we first investigate two types of noisy-representation contrastive learning: one applied between full-band discrete and band-split continuous representations, and another between band-split discrete and continuous representations. Additionally, we propose a noisy-clean representation contrastive learning method to enhance the model’s denoising ability. By enabling mutual reinforcement between contrastive and supervised losses during training, the proposed approach significantly improves performance without introducing extra parameters or computational cost during inference. Experiments on the VoiceBank+DEMAND dataset demonstrate that all the proposed methods outperform the original BSRNN across various evaluation metrics, highlighting more efficient and effective solutions in BSRNN-based speech enhancement tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Using Contrastive Learning for Improving BSRNN-Based Speech Enhancement

  • Yu Liao,
  • Li Li,
  • Haixin Guan,
  • Yanhua Long

摘要

BSRNN, a band-split RNN model designed for tasks such as speech enhancement and source separation, has demonstrated outstanding performance in recent research. Meanwhile, contrastive learning, as a self-supervised learning framework, has proven effective in guiding networks to learn speech representations with broad potential across domains. This paper explores integrating contrastive learning techniques into BSRNN to optimize its speech enhancement capabilities. Specifically, we design various contrastive losses that complement the supervised loss to jointly train the model. To improve the shallow acoustic representations generated by BSRNN’s band-split module, we first investigate two types of noisy-representation contrastive learning: one applied between full-band discrete and band-split continuous representations, and another between band-split discrete and continuous representations. Additionally, we propose a noisy-clean representation contrastive learning method to enhance the model’s denoising ability. By enabling mutual reinforcement between contrastive and supervised losses during training, the proposed approach significantly improves performance without introducing extra parameters or computational cost during inference. Experiments on the VoiceBank+DEMAND dataset demonstrate that all the proposed methods outperform the original BSRNN across various evaluation metrics, highlighting more efficient and effective solutions in BSRNN-based speech enhancement tasks.