This study introduces an advanced lip-reading technology that transforms video streams into textual content. It utilizes Spatio-Temporal Convolutional Neural Networks (ST-CNNs), Bidirectional-Recurrent Neural Networks (Bi-RNNs), and Connectionist Temporal Classification (CTC) to effectively calculate loss and capture the dynamics of visual speech. The ST-CNN processes both spatial and temporal aspects, while Bi-RNN layers interpret sequence dependencies. The integration of CTC loss helps optimize the training process by aligning video frames with corresponding text efficiently. This trainable system is highly adaptable and accurate, managing different speech patterns and environmental conditions with ease. Tests show significant improvements over traditional LSTM-based models. The model reports a Character Error Rate (CER) of 2.3% and Word Error Rate (WER) of 5.0%, showing substantial enhancements from the previous 11.6% WER with 2D models. These results confirm the model's advanced capabilities in recognizing visual speech. Extensive evaluations confirm its effectiveness, providing a robust solution for real-time video-to-text conversion. Furthermore, the development of this technology supports the goals of SDG 10, that aims to improve socio-economic and political inclusion for all, including those with hearing disabilities, by reducing communication barriers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Lip-Reading Techniques: Leveraging ST-CNN-BiGRU and CTC-Greedy Decoding for Improved Text Transcription

  • A. Kavya Varshini,
  • C. Jayaashri,
  • S. Padmini

摘要

This study introduces an advanced lip-reading technology that transforms video streams into textual content. It utilizes Spatio-Temporal Convolutional Neural Networks (ST-CNNs), Bidirectional-Recurrent Neural Networks (Bi-RNNs), and Connectionist Temporal Classification (CTC) to effectively calculate loss and capture the dynamics of visual speech. The ST-CNN processes both spatial and temporal aspects, while Bi-RNN layers interpret sequence dependencies. The integration of CTC loss helps optimize the training process by aligning video frames with corresponding text efficiently. This trainable system is highly adaptable and accurate, managing different speech patterns and environmental conditions with ease. Tests show significant improvements over traditional LSTM-based models. The model reports a Character Error Rate (CER) of 2.3% and Word Error Rate (WER) of 5.0%, showing substantial enhancements from the previous 11.6% WER with 2D models. These results confirm the model's advanced capabilities in recognizing visual speech. Extensive evaluations confirm its effectiveness, providing a robust solution for real-time video-to-text conversion. Furthermore, the development of this technology supports the goals of SDG 10, that aims to improve socio-economic and political inclusion for all, including those with hearing disabilities, by reducing communication barriers.