Context-Aware Enhanced Dereverberation Network
摘要
Reverberation is a common phenomenon in acoustic environments that significantly impairs the clarity and intelligibility of speech signals, consequently affecting the performance of various speech processing systems. To address the challenges of dynamic temporal importance and computational cost in existing deep learning dereverberation methods, this study proposes a novel dereverberation method called Context-Aware Enhanced Dereverberation Network (CAEDN). CAEDN employs a dual-stream architecture that integrates convolutional networks with self-attention mechanism, effectively modeling both short-term and long-term acoustic features in the spectrum. Additionally, CAEDN uses Mel spectrum as the feature representation, which is more consistent with human auditory perception and can improve dereverberation performance while reducing computational cost. The global self-attention module can dynamically focus on both temporal and frequency contexts, improving the accuracy and efficiency of the dereverberation model. Experiments on the WHAMR dataset show that CAEDN achieves state-of-the-art performance on three evaluation metrics: PESQ, SISNR, and ESTOI, demonstrating its robustness and effectiveness in different reverberant environments.