Progressive Retention Sampling for Sequence Generation-Based Scene Text Spotting
摘要
Sequence generation models have demonstrated promising results in scene text spotting. However, these models face a discrepancy between the training and inference phases: during training, ground-truth sequences are provided as input, whereas during inference, this input is replaced by the model’s own predictions. While current sampling strategies have effectively mitigated this discrepancy in various sequence generation tasks such as image captioning and machine translation, they are not directly applicable to scene text spotting due to its unique sequence structure. This paper introduces Progressive Retention Sampling, a novel sampling strategy tailored specifically for sequence generation-based scene text spotting. We evaluate our approach using two scene text spotting models, UNITS and SPTS, conducting experiments on the ICDAR 2015, Total-Text, and VinText datasets. Our results demonstrate that the proposed method outperforms both baselines and conventional sampling strategies.