Correction of textual errors in early-modern Japanese books
摘要
Recent, efforts to digitize books have aimed to enhance accessibility and convenience. However, early-modern Japanese books published between the Meiji and early Showa period (1868–1935) differ significantly from modern publications, posing challenges for character recognition systems. Although progress has been made in developing recognition systems tailored to these texts, recognition errors remain common. This paper proposes a method for detecting and correcting misrecognized characters in early-modern Japanese books using Bidirectional Encoder Representations from Transformers (BERT). The approach consists of three stages: erroneous sentence detection, recognition error identification, and error correction. Experimental results demonstrate that the proposed method achieves