We propose contour-guided context learning (CCL) for bilingual scene text recognition (STR). The CCL framework consists of three parts: Contour Guided Transformer (CGT), Contextual Learning Transformer (CLT) and Multimodal Transformer (MMT) for fusion. CGT embeds a CLIP image encoder and utilizes CLIP’s pre-training capabilities to capture contour features from input images, and CLT embeds a CLIP text encoder to correct contextual errors. The fusion network incorporates attention features extracted by Transformer to enhance text recognition performance. Unlike most STR methods that only target English, the proposed CCL is designed to handle both English and Chinese and can handle irregularly shaped scene text. We conduct a comprehensive evaluation on Chinese and English benchmark datasets to validate the performance of our approach against state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Contour-Guided Context Learning for Scene Text Recognition

  • Wei-Chun Hsieh,
  • Gee-Sern Hsu,
  • Jun-Yi Chen,
  • Moi Hoon Yap,
  • Zi-Chun Chao

摘要

We propose contour-guided context learning (CCL) for bilingual scene text recognition (STR). The CCL framework consists of three parts: Contour Guided Transformer (CGT), Contextual Learning Transformer (CLT) and Multimodal Transformer (MMT) for fusion. CGT embeds a CLIP image encoder and utilizes CLIP’s pre-training capabilities to capture contour features from input images, and CLT embeds a CLIP text encoder to correct contextual errors. The fusion network incorporates attention features extracted by Transformer to enhance text recognition performance. Unlike most STR methods that only target English, the proposed CCL is designed to handle both English and Chinese and can handle irregularly shaped scene text. We conduct a comprehensive evaluation on Chinese and English benchmark datasets to validate the performance of our approach against state-of-the-art methods.