<p>Handwritten optical character recognition (OCR) systems are often erroneous, owing to a combination of the variability in handwriting and technological limitations with respect to OCR. This paper describes a unique strategy for automatic OCR error correction based on both semantic similarity from WordNet and frequency-based rankings of words. The methodology identifies erroneous words, generates candidate corrections, and selects the best match, utilizing a combination of semantic and statistical analysis. The technique thus improves the correctness of correction without needing a large training set, and it is computationally efficient as well as adaptable to both document digitization and automated transcription. Proposed methods clearly show improved reliability in OCR texts while still upholding grammatical correctness and contextual accuracy. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive OCR Error Correction for Handwritten Texts: A Semantic and Statistical Approach

  • Sonakshi Vij,
  • Amita Jain,
  • Devendra Tayal,
  • Vishal Kumar,
  • Raunit Arora,
  • Ritwik Arora

摘要

Handwritten optical character recognition (OCR) systems are often erroneous, owing to a combination of the variability in handwriting and technological limitations with respect to OCR. This paper describes a unique strategy for automatic OCR error correction based on both semantic similarity from WordNet and frequency-based rankings of words. The methodology identifies erroneous words, generates candidate corrections, and selects the best match, utilizing a combination of semantic and statistical analysis. The technique thus improves the correctness of correction without needing a large training set, and it is computationally efficient as well as adaptable to both document digitization and automated transcription. Proposed methods clearly show improved reliability in OCR texts while still upholding grammatical correctness and contextual accuracy.