A new method for vocal activity detection (VAD) using a combination of CNN and BiLSTM architecture is presented in the abstract. The effective detection of speech in loud surroundings or with complicated audio patterns is sometimes a difficulty for traditional VAD systems. To improve VAD performance, the suggested hybrid model merges CNNs’ feature extraction capabilities with BiLSTMs’ temporal dependency capture capabilities. The convolutional neural network (CNN) starts by collecting both local and global patterns in the input audio spectrogram by extracting high-level features. After that, a BiLSTM network is trained with the features, and it uses the current time to improve its categorization. This setup enables the model to successfully distinguish between segments that include voice and those that do not, even when faced with difficult acoustic environments. Compared to more conventional approaches and independent architectures, the hybrid CNN-BiLSTM VAD performs better in experimental findings on benchmark datasets. When it comes to accuracy, resilience, and computing economy, the suggested model reaches state-of-the-art performance. In addition, it shows promise in generalizing to other types of audio and different speakers. The hybrid CNN-BiLSTM VAD is a huge step forward in the field of audio signal processing; it might be used for things like speaker diarization, audio surveillance, and speech recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Integration of Target Speaker Voice Activity Detection with Transformers and End-to-End Neural Networks

  • D. Anandha Kumar,
  • K. V. Narendra Nadh,
  • M. Sashank Chowdary

摘要

A new method for vocal activity detection (VAD) using a combination of CNN and BiLSTM architecture is presented in the abstract. The effective detection of speech in loud surroundings or with complicated audio patterns is sometimes a difficulty for traditional VAD systems. To improve VAD performance, the suggested hybrid model merges CNNs’ feature extraction capabilities with BiLSTMs’ temporal dependency capture capabilities. The convolutional neural network (CNN) starts by collecting both local and global patterns in the input audio spectrogram by extracting high-level features. After that, a BiLSTM network is trained with the features, and it uses the current time to improve its categorization. This setup enables the model to successfully distinguish between segments that include voice and those that do not, even when faced with difficult acoustic environments. Compared to more conventional approaches and independent architectures, the hybrid CNN-BiLSTM VAD performs better in experimental findings on benchmark datasets. When it comes to accuracy, resilience, and computing economy, the suggested model reaches state-of-the-art performance. In addition, it shows promise in generalizing to other types of audio and different speakers. The hybrid CNN-BiLSTM VAD is a huge step forward in the field of audio signal processing; it might be used for things like speaker diarization, audio surveillance, and speech recognition.