错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimization Algorithm of Visual Multimodal Text Recognition for Public Opinion Analysis Scenarios

  • Xing Liu,
  • Fupeng Wei,
  • Qiusheng Zheng,
  • Wei Jiang,
  • Liyue Niu,
  • Jizong Liu,
  • Shangshou Wang

摘要

Existing techniques for Monitoring public opinion on the internet generally rely on routinely mining text content from web pages, but they are unable to swiftly and effectively identify text content in images and videos. Therefore, a major challenge for multimodal information extraction in internet opinion scenarios is the quick and reliable identification and recognition of textual content in images and videos. Based on the state of the art in the field of optical character recognition (OCR), this paper proposed a improved method to visual multimodal text identification for scenarios such as internet opinion analysis. In order to enhance the learning effect of our overall model, this work upgrades the collaborative mutual learning (CML) distillation approach in the text detection module based on the combination of a classic distillation strategy and a deep mutual learning (DML) strategy. Next, the large kernel pixel aggregation network (LK-PAN) is then suggested as a solution to the earlier inadequacy in identifying multi-scale and text with extreme aspect ratios. In order to efficiently mine the contextual data in images or videos and to fulfill the goal of enhancing text recognition's mistake correcting capabilities, Transformer is finally implemented in the text recognition module. According to the experimental findings on the video dataset, our technique increases the F1 score by 17.97% and the recognition speed by 29.7%. The model provides important technical support for public opinion analysis in multimodal fields.