<p>Weak speech enhancement technology can improve the clarity and intelligibility of low-intensity speech in noisy environments, reduce the impact of background noise, and improve measurement accuracy. In this paper, we propose a joint time-frequency Cycle-Mixed Invariant Training (TF-Cycle-MixIT) method for unsupervised speech enhancement, specifically designed to address the challenge of enhancing weak speech signals in strong background noise. This approach integrates the MixIT network with the cyclic continuous learning method, overcoming the limitations of traditional speech separation models that heavily depend on pure speech signals as supervisory data. By fusing harmonic features with time-domain information, the model gains a deeper understanding of the intrinsic structure of speech signals, enabling precise denoising in complex acoustic environments. Field experiments conducted in a highway tunnel with strong background noise interference demonstrate the method’s effectiveness in recovering weak speech signals captured by a Distributed Acoustic Sensing system(DAS). The results show that our approach outperforms existing state-of-the-art (SOTA) unsupervised speech enhancement methods on the TIMIT dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Weak Speech Enhancement Using Periodic Mixing Invariant Training

  • Maoning Wang,
  • Xiaomin Bai,
  • Chensi Zhang,
  • Yuzhong Zhong

摘要

Weak speech enhancement technology can improve the clarity and intelligibility of low-intensity speech in noisy environments, reduce the impact of background noise, and improve measurement accuracy. In this paper, we propose a joint time-frequency Cycle-Mixed Invariant Training (TF-Cycle-MixIT) method for unsupervised speech enhancement, specifically designed to address the challenge of enhancing weak speech signals in strong background noise. This approach integrates the MixIT network with the cyclic continuous learning method, overcoming the limitations of traditional speech separation models that heavily depend on pure speech signals as supervisory data. By fusing harmonic features with time-domain information, the model gains a deeper understanding of the intrinsic structure of speech signals, enabling precise denoising in complex acoustic environments. Field experiments conducted in a highway tunnel with strong background noise interference demonstrate the method’s effectiveness in recovering weak speech signals captured by a Distributed Acoustic Sensing system(DAS). The results show that our approach outperforms existing state-of-the-art (SOTA) unsupervised speech enhancement methods on the TIMIT dataset.