TrOCR Meets Language Models: An End-to-End Post-correction Approach
摘要
This study aims to enhance handwritten text recognition (HTR) performance and domain adaptability by combining an optical character recognition (OCR) model with a language model (LM) that serves as a corrector. This integration addresses three principal challenges: over-correction, which compromises text authenticity; poor domain adaptation; and the scarcity of annotated images. We explore the synergy between TrOCR, a state-of-the-art OCR model, and CharBERT, a BERT-based LM. A novel aspect of our research involves introducing common errors made by the recogniser into the LM, enabling it to consider these errors during correction, thereby improving overall performance. Our findings reveal that the hybrid TrOCR-CharBERT model effectively balances visual and linguistic information, preserving the authenticity of the original texts. Furthermore, the model is able to adapt to historical data even when the recogniser is trained solely on contemporary data, mitigating the need for a large number of annotated historical handwritten images.