A Backend-Friendly On-Device Multi-channel Speech Enhancement System with IPD and PHM
摘要
Speech is the most direct way of communication between people, but noise reduces the clarity and intelligibility of speech signals. Although some mono-channel speech enhancement (SE) systems have been successfully applied to suppress noise, designing a small-footprint multi-channel SE system is still challenging, especially under extremely low signal-to-noise ratio (SNR). In this paper, we use interaural phase differences (IPD) and fixed-beam spatial information to construct a novel lightweight multi-channel SE system to ensure both performance improvement and a small footprint. First, we combine the widely used mono-channel enhancement model DCCRN with multi-channel directional features to make it applicable to multi-channel tasks by using IPD. Second, we improve the model to make it applicable to multi-channel tasks. Moreover, the information provided by the microphone array can significantly improve the performance of the SE system. Second, we compress the number of parameters and implement the Parameterized Hypercomplex Multiplication (PHM) method to meet the demands of limited memory resources. Finally, the experimental results show that our proposed system not only achieves better front-end performance in a low SNR environment but also boosts the backend performance, like the accuracy in keyword spotting systems, which indicates the superior denoising ability and less distortion of our proposed model.