<p>This paper introduces an Enhanced Chinese Scene Text Recognition (CSTR) Model leveraging Cross-Domain Feature Fusion to address the intricate challenges in CSTR, encompassing text deformation, vertical layout, complex structures, and an abundance of similar-looking characters. By incorporating frequency domain analysis and a detail enhancement mechanism, the model disentangles content and directional features in the frequency domain during visual feature extraction, markedly reducing confusion in recognizing similar-looking characters. The model’s architecture seamlessly integrates a ResNet encoder with a Transformer decoder, fusing spatial domain, frequency domain, and detailed features to bolster recognition performance. Comprehensive evaluations reveal that this model surpasses existing approaches, demonstrating superior accuracy and robustness in CSTR tasks, thereby advancing the state-of-the-art in Chinese scene text recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Chinese scene text recognition model base on cross-domain feature fusion

  • Ran Cui,
  • Aichun Zhu,
  • Zichen Ding

摘要

This paper introduces an Enhanced Chinese Scene Text Recognition (CSTR) Model leveraging Cross-Domain Feature Fusion to address the intricate challenges in CSTR, encompassing text deformation, vertical layout, complex structures, and an abundance of similar-looking characters. By incorporating frequency domain analysis and a detail enhancement mechanism, the model disentangles content and directional features in the frequency domain during visual feature extraction, markedly reducing confusion in recognizing similar-looking characters. The model’s architecture seamlessly integrates a ResNet encoder with a Transformer decoder, fusing spatial domain, frequency domain, and detailed features to bolster recognition performance. Comprehensive evaluations reveal that this model surpasses existing approaches, demonstrating superior accuracy and robustness in CSTR tasks, thereby advancing the state-of-the-art in Chinese scene text recognition.