Unsupervised Classification Model Based on Optimal Weighted Features for Voice Activity Detection
摘要
Voice activity detection (VAD) is an important task used in digital speech processing to identify active voice frames in audio signals. In this work, we propose a novel VAD method based on K-means classification using the optimal weighted features of audio frames. As a major contribution, we also carry out an online automatic adjustment of class-centroids to improve the VAD accuracy. The performances of the proposed model have been evaluated via normalized audio databases and real-time recorded signals in various noisy environments. The results show that the weighted K-mean clustering VAD acts robustly against moderate and strongly noisy situations, comparatively to existing unsupervised and recent adaptative thresholding systems.