Hybrid speech enhancement in modulation domain
摘要
This paper presents a comprehensive study on speech enhancement (SE) techniques, particularly focusing on the utilization of the discrete cosine transform (DCT) in the modulation domain (MD) in combination with the minimum mean square error (MMSE) and Kalman filter (KF) algorithms. The research investigates the efficacy of these techniques in improving speech quality and intelligibility under various noisy conditions, such as babble, drilling, horn, train, and traffic noises, across different signal-to-noise ratio (SNR) levels. Additionally, the study evaluates the performance of recurrent neural network (RNN) algorithms alongside traditional SE approaches. The study highlights the advantages of employing DCT in the modulation domain over the traditional discrete Fourier transform (DFT) in SE applications. Experimental results demonstrate significant improvements in objective quality metrics, including Perceptual Evaluation of Speech Quality (PESQ), composite measures, and short-time objective intelligibility (STOI), when implementing SE algorithms using discrete cosine transform in modulation domain. The proposed hybrid speech enhancement techniques, leveraging minimum mean square error and Kalman filter in combination with discrete cosine transform in modulation domain, outperform individual speech enhancement algorithms, showcasing superior noise reduction capabilities. This paper contributes to the advancement of speech enhancement methodologies, particularly in real-world noisy environments, and underscores the effectiveness of discrete cosine transform-based approaches in enhancing speech quality and intelligibility.