Enhancing Speech Quality Using Spectral Subtraction and Time-Frequency Filtering
摘要
Speech enhancement plays a vital role in almost all speech processing applications, where spectral subtraction based on voice activity detection (SS-VAD) is most commonly used. SS-VAD works well for moderate background noise conditions but deteriorates for low signal-to-noise-ratio (SNR) signal due to resulting residual noise consisting of musical tones. To address this issue, we introduce a novel approach called SS-time-frequency (SS-TF) filtering for enhancing speech quality. In this innovative method, we replace the conventional residual noise reduction technique with time-frequency (TF) filtering to effectively reduce additive noise. According to the perceptual evaluation of speech quality (PESQ), mean SNR and normalized covariance metric (NCM), the proposed approach outperforms SS-VAD for low SNR conditions. Further, estimation of average memory consumption and execution time has been done for both SS-VAD and the proposed methods.