<p>Automatic video retrieval and summarization are&#xa0;two of many applications that can be greatly aided by video subtitle recognition. However, because of complicated backgrounds, a wide variety of fonts, low contrast between words and backgrounds and lack of understanding of the text's context, existing algorithms face significant difficulties when it comes to video subtitle detection. The viewer may not receive subtitles that fully capture the essence of the video. Thus, the given paper presents an end-to-end pipeline approach for video context subtitle generation. The proposed hybrid approach takes input as raw video frames followed by Tesseract Optical Character Recognition (OCR) for text recognition. It is followed by the use of Bidirectional Encoder Representations from Transformers (BERT) and convolutional neural network (CNN) for contextual embedding of text and features extraction from video frames respectively. Lastly, long-short term memory (LSTM) takes input from CNN and BERT to understand temporal dependencies. The results show that the proposed hybrid approach yields 94.3% and 96.3% recognition accuracy on two publicly available datasets viz. TED Talks dataset and Open Images Dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced subtitle generation in videos: leveraging hybrid BERT–CNN–LSTM architecture for contextual understanding

  • Nagendra Babu Rajaboina,
  • Tulasi Prasad Sariki

摘要

Automatic video retrieval and summarization are two of many applications that can be greatly aided by video subtitle recognition. However, because of complicated backgrounds, a wide variety of fonts, low contrast between words and backgrounds and lack of understanding of the text's context, existing algorithms face significant difficulties when it comes to video subtitle detection. The viewer may not receive subtitles that fully capture the essence of the video. Thus, the given paper presents an end-to-end pipeline approach for video context subtitle generation. The proposed hybrid approach takes input as raw video frames followed by Tesseract Optical Character Recognition (OCR) for text recognition. It is followed by the use of Bidirectional Encoder Representations from Transformers (BERT) and convolutional neural network (CNN) for contextual embedding of text and features extraction from video frames respectively. Lastly, long-short term memory (LSTM) takes input from CNN and BERT to understand temporal dependencies. The results show that the proposed hybrid approach yields 94.3% and 96.3% recognition accuracy on two publicly available datasets viz. TED Talks dataset and Open Images Dataset.