<p>The primary challenge in speech enhancement lies in accurately tracking non-stationary noise over extended speech segments with low signal-to-noise ratios (SNR). Although numerous speech enhancement techniques have been developed over the past several decades, most remain ineffective in handling highly non-stationary noise. This study compares transform-based nonlinear speech enhancement techniques, emphasizing the spectral and temporal domains within a multiband framework. An oversampled wavelet packet filterbank is utilized to decompose noisy speech into seventeen subbands, while the noisy speech spectrum is partitioned into six unevenly frequency-spaced bands. The frequency distribution in both cases closely resembles the Bark scale separation in human auditory perception. Nonlinear spectral subtraction with a variable oversubtraction factor is applied to each frequency band to adapt to SNR variations across the speech spectrum. The performance of the proposed transform-based nonlinear frameworks is evaluated using both objective and subjective measures. Objective metrics include SNR, segmental SNR (SegSNR), Itakura–Saito Distance (ISD), and Perceptual Evaluation of Speech Quality (PESQ) scores under various stationary and non-stationary noise conditions and SNR levels. Subjective evaluation is conducted through Mean Opinion Score (MOS) ratings and spectrogram analysis. Experimental results demonstrate that the proposed temporal and spectral multiband framework achieve superior noise suppression and speech quality compared to conventional speech enhancement methods, as evidenced by improvements across multiple objective and subjective performance measures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transform-based nonlinear speech enhancement for monaural scenarios

  • Navneet Upadhyay,
  • Munir Georges

摘要

The primary challenge in speech enhancement lies in accurately tracking non-stationary noise over extended speech segments with low signal-to-noise ratios (SNR). Although numerous speech enhancement techniques have been developed over the past several decades, most remain ineffective in handling highly non-stationary noise. This study compares transform-based nonlinear speech enhancement techniques, emphasizing the spectral and temporal domains within a multiband framework. An oversampled wavelet packet filterbank is utilized to decompose noisy speech into seventeen subbands, while the noisy speech spectrum is partitioned into six unevenly frequency-spaced bands. The frequency distribution in both cases closely resembles the Bark scale separation in human auditory perception. Nonlinear spectral subtraction with a variable oversubtraction factor is applied to each frequency band to adapt to SNR variations across the speech spectrum. The performance of the proposed transform-based nonlinear frameworks is evaluated using both objective and subjective measures. Objective metrics include SNR, segmental SNR (SegSNR), Itakura–Saito Distance (ISD), and Perceptual Evaluation of Speech Quality (PESQ) scores under various stationary and non-stationary noise conditions and SNR levels. Subjective evaluation is conducted through Mean Opinion Score (MOS) ratings and spectrogram analysis. Experimental results demonstrate that the proposed temporal and spectral multiband framework achieve superior noise suppression and speech quality compared to conventional speech enhancement methods, as evidenced by improvements across multiple objective and subjective performance measures.