Shaky and non-shaky videos are quite common in real-time applications such as surveillance and monitoring vehicles and human movements in protected areas. As a result, text detection in such videos is a formidable challenge due to motion blur, noise, shaky cameras, poor quality and poor visibility. In contrast to existing text detection methods, which focus on text detection in scene images or specific types of images, the present work focuses on text detection in both shaky and non-shaky video frames. Inspired by the impressive performance of the HourGlass network for adverse situations, we explore the HourGlass network for successful text detection in shaky, non-shaky video frames and natural scene images. To improve the performance of the HourGlass network, we employ the Real-Time Model (RTMHead) for predicting text precisely and the Cross Stage Partial Network (CSPNet), which is a neck architecture for robust feature fusion. In addition, the integration of the SiLu activation function with the HourGlass network improves the discriminative power ability. To test the efficacy of the proposed method, we conducted experiments on shaky and non-shaky video frames, as well as ICDAR 2015 video frames. Furthermore, to show the effectiveness of the proposed method, we used Total-Text scene images for experimentation. The results on different datasets and a comparative study with the state-of-the-art models show that the proposed model outperforms the existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A New HourGlass Network for Detecting Text in Shaky and Non-shaky Video Frames

  • Arnab Halder,
  • Shivakumara Palaiahnakote,
  • Umapada Pal,
  • Michael Blumenstein,
  • Shivanand S. Gornale

摘要

Shaky and non-shaky videos are quite common in real-time applications such as surveillance and monitoring vehicles and human movements in protected areas. As a result, text detection in such videos is a formidable challenge due to motion blur, noise, shaky cameras, poor quality and poor visibility. In contrast to existing text detection methods, which focus on text detection in scene images or specific types of images, the present work focuses on text detection in both shaky and non-shaky video frames. Inspired by the impressive performance of the HourGlass network for adverse situations, we explore the HourGlass network for successful text detection in shaky, non-shaky video frames and natural scene images. To improve the performance of the HourGlass network, we employ the Real-Time Model (RTMHead) for predicting text precisely and the Cross Stage Partial Network (CSPNet), which is a neck architecture for robust feature fusion. In addition, the integration of the SiLu activation function with the HourGlass network improves the discriminative power ability. To test the efficacy of the proposed method, we conducted experiments on shaky and non-shaky video frames, as well as ICDAR 2015 video frames. Furthermore, to show the effectiveness of the proposed method, we used Total-Text scene images for experimentation. The results on different datasets and a comparative study with the state-of-the-art models show that the proposed model outperforms the existing methods.