<p>Source separation in highly reverberant environments is an open and extremely challenging issue in audio signal processing. High reverberation affects acoustic quality and deepens the difficulty of separating sound sources. In this paper, an equalizer network-based adaptive time-frequency source separation (EN-ATFSS) algorithm is proposed to deal with the source separation problem in highly reverberant environments, which can adaptively attenuate audible echoes and enhance the harmonic structure of real-world audio signals, achieving excellent separation performance. First, an equalizer network technology is designed to reshape the room impulse response and weaken audible echoes without changing the perceived timbre. Second, an adaptive Wiener-like mask is designed to enhance the dominant signals and attenuate the other components that have less energy than the dominant signals. Experimental results in public audio data sets demonstrate that the proposed EN-ATFSS algorithm improves signal intelligibility and achieves a much better separation performance than the state-of-the-art algorithms. In particular, it showed an improvement in the average source-to-distortion ratio (SDR), source-to-interference ratio (SIR), and source-to-artifacts ratio (SAR) of 4.8, 8.1, and 3.3 dB at a reverberation time of 900 ms, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Equalizer Network-based Adaptive Time-Frequency Source Separation in Highly Reverberant Environments

  • Yuan Xie,
  • Zhangchi Wei,
  • Tao Zou,
  • Yifei Sun,
  • Minghui Huang

摘要

Source separation in highly reverberant environments is an open and extremely challenging issue in audio signal processing. High reverberation affects acoustic quality and deepens the difficulty of separating sound sources. In this paper, an equalizer network-based adaptive time-frequency source separation (EN-ATFSS) algorithm is proposed to deal with the source separation problem in highly reverberant environments, which can adaptively attenuate audible echoes and enhance the harmonic structure of real-world audio signals, achieving excellent separation performance. First, an equalizer network technology is designed to reshape the room impulse response and weaken audible echoes without changing the perceived timbre. Second, an adaptive Wiener-like mask is designed to enhance the dominant signals and attenuate the other components that have less energy than the dominant signals. Experimental results in public audio data sets demonstrate that the proposed EN-ATFSS algorithm improves signal intelligibility and achieves a much better separation performance than the state-of-the-art algorithms. In particular, it showed an improvement in the average source-to-distortion ratio (SDR), source-to-interference ratio (SIR), and source-to-artifacts ratio (SAR) of 4.8, 8.1, and 3.3 dB at a reverberation time of 900 ms, respectively.