<p>Voice activity detection (VAD) is an important task used in digital speech processing to identify active voice frames in audio signals. In this work, we propose a novel VAD method based on K-means classification using the optimal weighted features of audio frames. As a major contribution, we also carry out an online automatic adjustment of class-centroids to improve the VAD accuracy. The performances of the proposed model have been evaluated via normalized audio databases and real-time recorded signals in various noisy environments. The results show that the weighted K-mean clustering VAD acts robustly against moderate and strongly noisy situations, comparatively to existing unsupervised and recent adaptative thresholding systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Classification Model Based on Optimal Weighted Features for Voice Activity Detection

  • Rim Boukri,
  • Atef Farrouki

摘要

Voice activity detection (VAD) is an important task used in digital speech processing to identify active voice frames in audio signals. In this work, we propose a novel VAD method based on K-means classification using the optimal weighted features of audio frames. As a major contribution, we also carry out an online automatic adjustment of class-centroids to improve the VAD accuracy. The performances of the proposed model have been evaluated via normalized audio databases and real-time recorded signals in various noisy environments. The results show that the weighted K-mean clustering VAD acts robustly against moderate and strongly noisy situations, comparatively to existing unsupervised and recent adaptative thresholding systems.