Enhancing Voice Activity Detection in Noisy Environments Using Deep Neural Networks
摘要
Voice activity detection (VAD) is a crucial component in numerous speech processing applications. One of the primary challenges in VAD is achieving a balance between avoiding false negatives and minimizing false positives in noisy environments. Even the past works have reported promising accuracy, nevertheless the performance of VAD in negative signal-to-noise ratio (SNR) conditions remains an open question. In this work, we aim to establish a deep neural network (DNN)-based speech enhancement pre-processing framework to improve VAD accuracy under deep noisy conditions. Additionally, we explore a DNN-based techniques that leverage the phase component to enhance speech quality. Experimental results on NOIZEUS database show that our proposed approach outperforms vector quantization-based VAD (VQ-VAD), as demonstrated by the VAD metrics.