<p>Scene text recognition methods are broadly categorized into serial and parallel. Serial methods achieve superior accuracy but are slower in speed. Parallel methods offer faster speed but may sacrifice accuracy. Current methods struggle to strike a balance between accuracy and inference speed, particularly facing challenges in both accuracy and speed. Therefore, we propose a new scene text recognizer called SNFR. It includes a simple yet efficient decoder, Salient Neighbor Decoder (SND), which achieves high accuracy recognition with lower computational cost for attention map calculation. SND generates a neighbor matrix by selecting salient positions, which guides the generation of all the character attention maps. We also propose a Text Feature Refining Module (TFRM) to capture the contextual relationship of text sequences, enhancing the overall feature representation of scene text. The experimental results demonstrate that our method achieves competitive performance on standard datasets and also shows superior performance on long text recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SNFR: salient neighbor decoding and text feature refining for scene text recognition

  • Tongwei Lu,
  • Huageng Fan,
  • Yuqian Chen,
  • Pengyan Shao

摘要

Scene text recognition methods are broadly categorized into serial and parallel. Serial methods achieve superior accuracy but are slower in speed. Parallel methods offer faster speed but may sacrifice accuracy. Current methods struggle to strike a balance between accuracy and inference speed, particularly facing challenges in both accuracy and speed. Therefore, we propose a new scene text recognizer called SNFR. It includes a simple yet efficient decoder, Salient Neighbor Decoder (SND), which achieves high accuracy recognition with lower computational cost for attention map calculation. SND generates a neighbor matrix by selecting salient positions, which guides the generation of all the character attention maps. We also propose a Text Feature Refining Module (TFRM) to capture the contextual relationship of text sequences, enhancing the overall feature representation of scene text. The experimental results demonstrate that our method achieves competitive performance on standard datasets and also shows superior performance on long text recognition.