Integrated noise suppression techniques for enhancing voice activity detection in degraded environments
摘要
Voice activity detection (VAD) is a crucial task in many speech processing applications, particularly in environments with low signal-to-noise ratios (SNR), where distinguishing speech from background noise is challenging. Deep learning-based techniques have achieved better performance compared to traditional signal processing-based techniques for real-time speech processing applications. However, most deep learning-based approaches require large amounts of training data and may not be suitable for resource-constrained (limited data) environments. In this paper, a noise suppression block is integrated to improve the performance of VAD in various real-world acoustic scenarios. Initially, we assess the soft mask estimator with a priori SNR uncertainty (SMPR) and soft mask estimator with posteriori SNR uncertainty (SMPO) techniques using speech quality metrics. Subsequently, the enhanced speech signal is applied for detecting voice activity. Experimental results on the NOIZEUS dataset for various noise types and levels show that the proposed method can improve VAD performance.