Sequence generation models have demonstrated promising results in scene text spotting. However, these models face a discrepancy between the training and inference phases: during training, ground-truth sequences are provided as input, whereas during inference, this input is replaced by the model’s own predictions. While current sampling strategies have effectively mitigated this discrepancy in various sequence generation tasks such as image captioning and machine translation, they are not directly applicable to scene text spotting due to its unique sequence structure. This paper introduces Progressive Retention Sampling, a novel sampling strategy tailored specifically for sequence generation-based scene text spotting. We evaluate our approach using two scene text spotting models, UNITS and SPTS, conducting experiments on the ICDAR 2015, Total-Text, and VinText datasets. Our results demonstrate that the proposed method outperforms both baselines and conventional sampling strategies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Progressive Retention Sampling for Sequence Generation-Based Scene Text Spotting

  • Phuc Nguyen,
  • Van-Manh Thai,
  • Xuan-Nguyen Pham,
  • Trong-Le Do

摘要

Sequence generation models have demonstrated promising results in scene text spotting. However, these models face a discrepancy between the training and inference phases: during training, ground-truth sequences are provided as input, whereas during inference, this input is replaced by the model’s own predictions. While current sampling strategies have effectively mitigated this discrepancy in various sequence generation tasks such as image captioning and machine translation, they are not directly applicable to scene text spotting due to its unique sequence structure. This paper introduces Progressive Retention Sampling, a novel sampling strategy tailored specifically for sequence generation-based scene text spotting. We evaluate our approach using two scene text spotting models, UNITS and SPTS, conducting experiments on the ICDAR 2015, Total-Text, and VinText datasets. Our results demonstrate that the proposed method outperforms both baselines and conventional sampling strategies.