TSCME-Net: Two-Stream Conformer-Enhanced Network Leveraging Complex and Magnitude Spectrum Modeling for Noise-Robust Speech Enhancement
摘要
Traditional methods for speech enhancement primarily focus on restoring amplitude features while neglecting phase information, which is equally critical for perceived quality. To address this issue, this paper proposes a dual-stream Conformer model that jointly processes complex and amplitude-domain features. The amplitude stream employs a mask to extract amplitude information, whereas the complex stream is responsible for capturing phase features. The model integrates Temporal Attention (TA), Dilated Convolution (DC), and Frequency Attention (FA) to extract both local and global speech features, with an attention-aware feature fusion (AFF) module for efficient dual-stream fusion, thereby enabling precise spectral estimation. The experimental results demonstrate that the proposed model significantly outperforms existing benchmarks on the VoiceBank + DEMAND dataset. Furthermore, ablation studies on the TIMIT dataset have been conducted to verify the contribution of each sub-module. The experimental results demonstrate that the proposed method achieves a 3.2% improvement over the latest methods on the dataset.