Adaptive OCR Error Correction for Handwritten Texts: A Semantic and Statistical Approach
摘要
Handwritten optical character recognition (OCR) systems are often erroneous, owing to a combination of the variability in handwriting and technological limitations with respect to OCR. This paper describes a unique strategy for automatic OCR error correction based on both semantic similarity from WordNet and frequency-based rankings of words. The methodology identifies erroneous words, generates candidate corrections, and selects the best match, utilizing a combination of semantic and statistical analysis. The technique thus improves the correctness of correction without needing a large training set, and it is computationally efficient as well as adaptable to both document digitization and automated transcription. Proposed methods clearly show improved reliability in OCR texts while still upholding grammatical correctness and contextual accuracy.