<p>Deep Neural Networks (DNNs) have been central to acoustic modeling for more than two decades. Convolutional Neural Networks (CNNs), an advancement over DNNs, effectively capture spectral variations and local correlations in speech signals. Similarly, Bidirectional Long Short-Term Memory (BLSTM) networks strengthen temporal modeling by learning higher-level representations, improving recognition accuracy in ASR systems. Moreover, most neural networks utilize the softmax activation function for prediction purposes. Therefore, this softmax layer may be replaced with Support Vector Machines (SVMs) to handle high-dimensional features more efficiently. Recognizing the strengths of these architectures, this study proposes a hybrid end-to-end model integrating CNN, BLSTM, and SVM modules. In this framework, CNNs perform feature extraction, BLSTMs capture temporal dependencies, and SVMs enhance classification in high-dimensional spaces. Instead of training each component individually, the parameters of hybrid architecture are together trained in an End-to-End manner. The proposed architecture also addresses CNNs’ limitation in modeling tonal features by introducing a fusion mechanism for tonal and raw features. Experiments explore the optimal BLSTM configuration and evaluate system robustness under noisy conditions. The hybrid CNN–BLSTM–SVM model achieves a 10.2% performance improvement over the baseline CNN and a 2.73% gain over the CNN–BLSTM system, demonstrating superior accuracy and resilience in challenging acoustic environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid CNN-BLSTM support vector machine architecture for E2E speech recognition

  • Jaspreet Kaur Sandhu,
  • Munish Kumar,
  • Amitoj Singh

摘要

Deep Neural Networks (DNNs) have been central to acoustic modeling for more than two decades. Convolutional Neural Networks (CNNs), an advancement over DNNs, effectively capture spectral variations and local correlations in speech signals. Similarly, Bidirectional Long Short-Term Memory (BLSTM) networks strengthen temporal modeling by learning higher-level representations, improving recognition accuracy in ASR systems. Moreover, most neural networks utilize the softmax activation function for prediction purposes. Therefore, this softmax layer may be replaced with Support Vector Machines (SVMs) to handle high-dimensional features more efficiently. Recognizing the strengths of these architectures, this study proposes a hybrid end-to-end model integrating CNN, BLSTM, and SVM modules. In this framework, CNNs perform feature extraction, BLSTMs capture temporal dependencies, and SVMs enhance classification in high-dimensional spaces. Instead of training each component individually, the parameters of hybrid architecture are together trained in an End-to-End manner. The proposed architecture also addresses CNNs’ limitation in modeling tonal features by introducing a fusion mechanism for tonal and raw features. Experiments explore the optimal BLSTM configuration and evaluate system robustness under noisy conditions. The hybrid CNN–BLSTM–SVM model achieves a 10.2% performance improvement over the baseline CNN and a 2.73% gain over the CNN–BLSTM system, demonstrating superior accuracy and resilience in challenging acoustic environments.