错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech Enhancement Using U-Net-Based Progressive Learning with Squeeze-TCN

  • Sunny Dayal Vanambathina,
  • Sivaprasad Nandyala,
  • Chaitanya Jannu,
  • J. Sirisha Devi,
  • Sivaramakrishna Yechuri,
  • Veeraswamy Parisae

摘要

Speech enhancement is crucial in many speech processing applications. Recently, researchers have been exploring ways to improve performance by effectively capturing the long-term contextual relationships within speech signals. Nevertheless, these investigations often neglect the significance of the time-frequency (T-F) distribution details of speech spectral elements, which are also important for improving speech quality. Moreover, accurately mapping spectral information in supervised methods requires both global and local information. This paper presents a novel speech enhancement model that utilizes long-term and time-frequency (T-F) information of speech spectral elements for enhancing target speakers. The system is composed of a feature extractor and reconstructor. The feature extractor introduces the U-Net block to enable feature recalibration from multiple scales. Furthermore, the time-frequency attention (TFA) within the U-net’s feature extractor delineates the prominent time-frequency distribution of speech through the simultaneous utilization of two branches: one for time-frame attention and another for frequency channel attention. TFA involves the assignment of distinct attention weights to individual time-frequency spectral components, empowering the models to concentrate on both the “when” (time frames) and the “what” (frequency-wise channels). In the reconstructor, several filtering modules (FMs) are stacked to facilitate the reconstruction of a spectrum in a progressive manner to obtain a coarse estimation by reducing noise in the magnitude domain. Furthermore, the intermediate outcome can be progressively improved across stages by unfolding the filtering modules repeatedly, ultimately resulting in the estimation of the spectrum. The proposed approach is tested on the Librispeech and Voicebank datasets. The findings indicate that this method surpasses earlier models and attains the state-of-the-art performance on these datasets.