Optical Character Recognition (OCR) is a revolutionary technique that aids machines in retrieving textual content from images to perform further analysis. However, OCR has its limitations, especially when dealing with degraded or low-quality images, which can impact the overall reliability of the text recognition process. Thus, the system’s accuracy is contingent upon the quality of the input (digital or handwritten documents). Efforts to modify the text detection and text recognition modules in existing OCRs fail to work in complex dynamic environments due to the complexity of the background information of the input data. Thus, a new first-of-its-kind annotated dataset called OCR-SBT for digital text segmentation is proposed in this work, along with a novel pre-processing pipeline using deep learning that performs text retrieval from images having varying and complex backgrounds using binary semantic segmentation. With quantitative metrics such as the DICE coefficient as high as 99.56%, the qualitative performance improvement of OCR has also been validated on real-world test samples containing varying contextual information to validate the model’s efficacy. Ablation experiments are also performed to determine the importance of super-resolution of input images using Stable Diffusion and ESRGAN. This work will help the research community to improve OCR for several real-world applications by alleviating the problems related to background contextual information obfuscating the text recognition module. The dataset and the codes will be made publically available at the Github Link: https://github.com/argon125/OCR-SBT-Performing-Text-Segmentation-to-Improve-OCR .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performing Text Segmentation to Improve OCR on Multi Scene Text

  • Arrun Sivasubramanian,
  • Sheel Shah,
  • Akash Narayanaswamy,
  • C. Rindhya,
  • H. B. Barathi Ganesh

摘要

Optical Character Recognition (OCR) is a revolutionary technique that aids machines in retrieving textual content from images to perform further analysis. However, OCR has its limitations, especially when dealing with degraded or low-quality images, which can impact the overall reliability of the text recognition process. Thus, the system’s accuracy is contingent upon the quality of the input (digital or handwritten documents). Efforts to modify the text detection and text recognition modules in existing OCRs fail to work in complex dynamic environments due to the complexity of the background information of the input data. Thus, a new first-of-its-kind annotated dataset called OCR-SBT for digital text segmentation is proposed in this work, along with a novel pre-processing pipeline using deep learning that performs text retrieval from images having varying and complex backgrounds using binary semantic segmentation. With quantitative metrics such as the DICE coefficient as high as 99.56%, the qualitative performance improvement of OCR has also been validated on real-world test samples containing varying contextual information to validate the model’s efficacy. Ablation experiments are also performed to determine the importance of super-resolution of input images using Stable Diffusion and ESRGAN. This work will help the research community to improve OCR for several real-world applications by alleviating the problems related to background contextual information obfuscating the text recognition module. The dataset and the codes will be made publically available at the Github Link: https://github.com/argon125/OCR-SBT-Performing-Text-Segmentation-to-Improve-OCR .