<p>Speech signals play a vital role in daily communication, professional tasks, and academic research. Nevertheless, in real-world situations, the target speech is frequently compromised by environmental noise, failing to meet human auditory expectations. Consequently, speech enhancement is required to reduce noise interference while minimizing distortion of the target speech. Since polynomial matrices can fully utilize the temporal, spatial, and frequency correlations of speech, they are often used in filter design. This paper proposes a speech enhancement method based on grouped polynomial eigenvalue decomposition (PEVD) method. Compared with conventional PEVD methods, the proposed approach employs a staged processing scheme. In the first stage, the microphone array observation signals are divided into groups, each containing two-channel signals. PEVD operations are performed on each group, and the resulting signal subspaces are used to filter the corresponding grouped signals, producing the first-stage output signals. These output signals then undergo the same process in the second stage: they are regrouped, processed with PEVD to obtain new signal subspaces, and filtered again. This grouping and decomposition process continues iteratively until further grouping is no longer feasible. Finally, all signal subspaces obtained across the different stages are combined to form the complete signal subspace representation. Simulations with linear microphone arrays in single-source scenarios demonstrate that, in terms of segmental signal to noise ratio (SegSNR), perceptual evaluation of speech quality (PESQ), and short-time objective intelligibility (STOI), the proposed method achieves comparable performance to traditional PEVD-based speech enhancement while reducing computational complexity by an average of 33.4%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech Enhancement Based on Grouped Polynomial Eigenvalue Decomposition

  • Yue Zhao,
  • Liang Tao,
  • Maoshen Jia,
  • Enchang Sun

摘要

Speech signals play a vital role in daily communication, professional tasks, and academic research. Nevertheless, in real-world situations, the target speech is frequently compromised by environmental noise, failing to meet human auditory expectations. Consequently, speech enhancement is required to reduce noise interference while minimizing distortion of the target speech. Since polynomial matrices can fully utilize the temporal, spatial, and frequency correlations of speech, they are often used in filter design. This paper proposes a speech enhancement method based on grouped polynomial eigenvalue decomposition (PEVD) method. Compared with conventional PEVD methods, the proposed approach employs a staged processing scheme. In the first stage, the microphone array observation signals are divided into groups, each containing two-channel signals. PEVD operations are performed on each group, and the resulting signal subspaces are used to filter the corresponding grouped signals, producing the first-stage output signals. These output signals then undergo the same process in the second stage: they are regrouped, processed with PEVD to obtain new signal subspaces, and filtered again. This grouping and decomposition process continues iteratively until further grouping is no longer feasible. Finally, all signal subspaces obtained across the different stages are combined to form the complete signal subspace representation. Simulations with linear microphone arrays in single-source scenarios demonstrate that, in terms of segmental signal to noise ratio (SegSNR), perceptual evaluation of speech quality (PESQ), and short-time objective intelligibility (STOI), the proposed method achieves comparable performance to traditional PEVD-based speech enhancement while reducing computational complexity by an average of 33.4%.