Improving low-complexity and real-time DeepFilterNet2 for personalized speech enhancement
摘要
DeepFilterNet2, a recently proposed real-time and low-complexity speech enhancement (SE) technique, has shown state-of-the-art SE performance in many deep noise suppression tasks. This paper proposes a new approach, termed pDeepFilterNet2, to generalize and improve the original DeepFilterNet2 for personalized speech enhancement (PSE) tasks under multi-talker noisy and reverberant conditions. Three improvements are investigated: a frame-wise speaker adaptation (FSA) is first proposed to achieve dynamic target speaker cues for generalizing the DeepFilterNet2 to pDeepFilterNet2; Then, we remove a redundant skip connection in original DeepFilterNet2 and add a simple layer of causal multi-head self-attention to enhance the model for aggregating global context information; Finally, a multi-domain loss function combing both time and frequency domain losses is introduced to further improve the PSE system performance. Moreover, different types of target speaker embedding are also investigated in this study to see their effectiveness. Our experiments are conducted on the DNS4 Challenge dataset and results show that the proposed pDeepFilterNet2 outperforms the original DeepFilterNet2 significantly across multiple evaluation metrics. Furthermore, it exhibits competitive performance when compared to our previously proposed sDPCCN, which was specifically designed for target speaker extraction.